Runtime Analysis · Ruby
Your agent guesses.
Your runtime knows.
It records what your objects actually do at runtime, and tells your coding agent and your editor.
Free with an account. Records your test suite on your laptop, with no production needed. Your recordings stay on your machine: sending anything to us is opt-in. A hosted tier for teams is coming.
The problem
The most important facts about your code aren't in your code.
What a method actually returns. Which collaborators an object really needs. Which of a class's methods each caller uses, and which seams your tests cross for real. Ruby writes none of it down: types say what someone declared, coverage says which lines ran, and grep finds text. So everyone fills the gap by guessing, and each guess goes wrong in a different place.
One recording of what actually runs answers all of it, starting with your test suite.
What runtime_analysis is
A map of what your code actually does.
It records the class of every object your code touches and every message it receives, builds a graph from that, and answers questions from the graph wherever you work.
In your process. The gem runs inside your test suite and records every message sent to every object: the receiver's class, the method called, the call site, in order. Class names and method names, never your data, written to offline files. Point it at your app in development, staging or production for real traffic; nothing needs it.
On your machine, or hosted. The messages sent to an object are the shape the code depends on. Follow a return value to where it's used next, and the methods called on it there are its type: duck typing, read from real runs. Locally the graph is built from your recorded files. With Pro, once you opt in, every deploy's runs pool into one hosted graph.
In your tools. Your agent over MCP, your editor over LSP, the command line and CI. They only read the graph; nothing they do touches your app.
Ask it anything
Every method, annotated with what actually ran through it.
Grep for .call and you get ten thousand hits: every service object, proc and
middleware in your app and its gems, with nothing to say which ones this line reaches. The
recording has the answer: the receivers that actually showed up here, with their types and
how often. Each tool is one question, asked of the same recording. A few of them:
Testing
- Does this stub match the real thing?verify_mock
- Which seams do my tests only cross through a double?testing_pyramid
- Which specs reach the method I just changed?specs_that_reach
- How do I build a real Invoice for this test?how_to_reach
- Is this unit test really a unit test?how_far_it_reached
- Which untested methods matter most?unit_test_candidates
Understanding
- What does this method actually return?observed_shape
- What does it get passed?what_it_receives
- Who calls it, and how often?what_calls
- How did execution get here?how_it_got_here
- Which classes actually answer this message?who_actually_answered
- Where does this nil come from?where_nil_comes_from
Changing
- What will changing this signature break?what_a_change_touches
- Is anything still calling this?dead_api
- Did this PR break a contract?check_diff
Performance
- Where does the time go?where_time_goes
- Is there an N+1 in this request?whats_looping
Design
- Which roles is this class playing?split_interface
- Does this method belong on another object?feature_envy
- Which file belongs in another folder?relocate_file
- Which dependency cycle should I cut, and where?break_cycle
Anything else? Ask the hosted graph directly in Cypher. See all the tools →
Examples
Three questions, followed all the way through.
Each starts from one question in that list and shows the answer where you'd actually use it.
Example 1 · a spec, written first
Same spec, written first. One of them has watched the collaborator run.
✗ agent alone · green, but wrong
1
2
3
4
5
6
7
8
# agent: failing spec first, stubbing the curriculum RSpec.describe Enrollment do it "unlocks the next book" do curriculum = instance_double(Curriculum) allow(curriculum).to receive(:unlocked_for).and_return([:geometry]) expect(Enrollment.new(curriculum:).unlock_next).to eq([:geometry]) end end
Curriculum#unlocked_for returns an Array whose elements receive #unlocked? and #id, in all 12 examples of curriculum_spec.rb and 3,181 production runs. :geometry is a Symbol; it responds to neither.
It passes green. The stub returns a value the real collaborator never would.
> Enrollment needs #unlock_next. Failing spec first, then make it pass
● I'll stub the curriculum and write the failing spec.
● Bash(bundle exec rspec spec/models/enrollment_spec.rb)
⎿ 1 example, 1 failure · undefined method 'unlock_next'
● Edit(app/models/enrollment.rb)
⎿ 1 example, 0 failures
● Red, then green. Done.
✓ agent + runtime analysis · fixed before it ran
1
2
3
4
5
6
7
8
9
# agent → runtime-analysis · observed_shape(Curriculum#unlocked_for) # returns Array of {#unlocked?, #id} RSpec.describe Enrollment do it "unlocks the next book" do book = instance_double(Book, unlocked?: true, id: 1) curriculum = instance_double(Curriculum, unlocked_for: [book]) expect(Enrollment.new(curriculum:).unlock_next).to eq([book]) end end
● The runtime-analysis hook flagged my stub: the double can't stand in for the real curriculum.
● runtime-analysis · observed_shape(Curriculum#unlocked_for) [MCP]
⎿ elements respond to #unlocked? #id · Book
● Rebuilding with a verified double that matches.
● Bash(bundle exec rspec spec/models/enrollment_spec.rb)
⎿ 1 example, 0 failures
● Tests pass, and now for the right reason.
The agent didn't guess the second time. It asked runtime-analysis for the real shape of
Curriculum#unlocked_for, recorded while Curriculum's own spec ran, and
built a double that matches it. Production runs add weight when you record them, but the spec
alone is enough, and no code was written ahead of its test. It passes because the code is right.
Example 2 · what doubles hide
Every unit green. The seam never ran for real.
A lying stub is one double. Zoom out and the same blind spot covers whole boundaries, test first
or test after. Checkout's spec stubs the gateway and the gateway's spec tests itself:
every line is covered, and no example ever runs the real Checkout calling the real
PaymentGateway. Coverage counts lines, not collaborations.
Runtime Analysis maps the recorded call graph onto your pyramid and names the boundaries only ever crossed through a double. Record production too and they're ranked by real traffic. Then it scaffolds the test that crosses one for real.
The same gap, on whichever surface you're working: your agent, your editor, the command line.
testing_pyramid(stubbed_only: true) → { stubbed_only: 212, boundaries: [ { call: "Checkout#call", to: "PaymentGateway#charge", stubbed: 3, real: 0, prod: 8204 }, { call: "Report#build", to: "Invoice#total", stubbed: 2, real: 0, prod: 4102 } ] }
1
2
3
def call gateway.charge(amount) ⤷ stubbed in 3 specs · never real end
$ ra testing_pyramid --stubbed-only 212 boundaries · only crossed via doubles Checkout#call → PaymentGateway#charge 3 stubbed · 0 real · 8,204 prod Report#build → Invoice#total 2 stubbed · 0 real · 4,102 prod
Then it writes the test that crosses the seam, from a real run: --scaffold on the CLI, the editor's
quick-fix, or your agent generating it from the structured boundary.
$ ra testing_pyramid --scaffold Checkout#call # setup taken from a real run (how_to_reach) RSpec.describe Checkout do it "charges the gateway" do gateway = PaymentGateway.new(account:, rate_table:) checkout = Checkout.new(cart:, gateway:) expect { checkout.call }.to change(gateway, :charged?).to(true) end end
Example 3 · an interface nobody designed
Because we track the whole interface, we know the part each caller uses.
A fat interface shows up when each caller uses a different, small part of a type, measured from real runs. It appears in your editor as a hint with a quick-fix, never a build error. Partial use is normal, so failing a build on it would only add noise.
1
2
3
4
5
6
7
8
class Curriculum def next_after(book) = sequence[book] def first_books = @first_books def unlocked_for(student) = eligible(student) def weeks_for(book) = @weeks[book] def final_books = @final_books end
Across your suite, Curriculum's callers split into 3 groups. No caller uses methods from more than one:
- Scheduler → { next_after, first_books }
- Enrollment → { unlocked_for }
- Report → { weeks_for, final_books }
💡 Quick Fix · Extract 3 role interfaces
Fat interfaces are one of many design questions the same graph answers: SOLID, the Gang-of-Four patterns, and package cohesion and coupling, each turned into a concrete move.
Types you observe, not declare
Sorbet and RBS make you name every interface. Your code already relies on implicit ones, and we read those off real runs, polymorphism and all.
See more →Concrete moves, measured
Feature envy, dependency cycles, unstable dependencies and packages used in parts, each with the move to make and the calls behind it.
See more →Where you ask
In your agent, your editor, your terminal and CI.
Curriculum#unlocked_for from one recording: queried by your agent, shown in your
editor, printed on the command line, and enforced in CI.
observed_shape( "Curriculum#unlocked_for", as: "return" ) → { return: { class: Array, responds: [select, each, any?] }, element: { class: Book, responds: [unlocked?, id, title] }, examples: 12, prod_runs: 3181 }
1
2
3
def unlock_next = curriculum.unlocked_for(student) ⤷ Array<Book> · #unlocked? #id #title
$ ra observed_shape \ Curriculum#unlocked_for return → #select #each #any? Array element → #unlocked? #id #title Book · from 12 spec examples and 3,181 production runs
▸ ra check --since origin/main ✗ verify_mock · Curriculum#unlocked_for stub returns Array<Symbol> real runs Array<Book> → #unlocked? #id #title ↳ 1 false positive introduced Error: Process completed with exit code 1 ↳ PRO posted to the PR as a review comment
An MCP tool your agent can call, plus a Claude Code hook that checks every spec it writes. A lying mock gets caught in the same edit that introduced it.
> Enrollment needs #unlock_next. Test first, then make it pass
● I'll stub the curriculum collaborator and write the failing spec.
● Write(spec/models/enrollment_spec.rb)
⎿ +12 lines
✗ verify_mock: your stub for Curriculum#unlocked_for returns Array<Symbol>; in curriculum_spec.rb and 3,181 production runs the real return receives #unlocked? and #id (Array<Book>). A Symbol responds to neither.
● The hook caught it before the spec even ran: my double can't stand in for the real curriculum. Let me pull the real shape.
● runtime-analysis · observed_shape(Curriculum#unlocked_for) [MCP]
⎿ return → #select #each · elements → #unlocked? #id #title · never nil
● Rebuilding the spec against the real contract.
● Write(spec/models/enrollment_spec.rb)
● Bash(bundle exec rspec …)
⎿ 1 example, 1 failure · undefined method 'unlock_next'
● Red, for the right reason. Now the implementation.
● Edit(app/models/enrollment.rb)
● Bash(rspec && mutant run …)
⎿ 1 example, 0 failures · mutations 24/24 killed
● Green, and every mutant of it dies.
What the agent runs over MCP is a plain command too. Run ra verify_mock in your
suite, or ra check on a diff as a CI gate, and a lying mock fails the build instead
of merging.
$ ra verify_mock enrollment_spec.rb ✗ stub Curriculum#unlocked_for is a false positive your stub returns [:geometry] Array<Symbol> the real one returns an Array whose elements receive #unlocked?, #id in curriculum_spec.rb (12 examples) and 3,181 production runs a Symbol responds to neither ↳ PRO real types responding to {#unlocked?, #id}: Book · SampleBook · ArchivedBook across your app → swap one in, or generate a verified fake from its contract ↳ and the bug it hid: unlock_next returned :geometry for a student who hasn't finished its prerequisite — a locked book leaked
However you test
Test first, test hard, or not yet.
The recorder runs wherever you point it: at a suite you write first, at a suite that already bites, or at an app with no tests yet.
Your suite is the recording
The collaborators you stub already exist, and each one is real in its own spec. No production, and no code written ahead of its test.
See more →What mutation testing can't see
For suites that already bite: check the doubles mutant never touches, aim it where it matters, and wire real collaborators without the tax.
See more →Record the app instead
Code with no specs still shows up once your app runs. It describes the real contract and writes the first test from an actual run.
See more →Test first
You write the test first. Good: your suite is the recording.
The class you're about to write doesn't exist yet, but everything it collaborates with does, and
earlier cycles' tests already ran that code for real. Mock Gateway in
Checkout's spec and Gateway's own spec still runs the real thing, so
every double has a real recording to be checked against.
The stub you just wrote is checked against what the real collaborator did in its own spec.
Run only the specs that reach what you changed, so the loop stays in seconds.
See the slice of the interface each caller uses, and the methods nothing but their own spec calls.
Test hard
Mutant proves your tests bite. It can't prove your mocks don't lie.
Run mutant and it mutates the code under test. It does not touch the double you stubbed
for a collaborator. instance_double checks the method exists with the right arity, not
the shape of what the real object returns. Runtime Analysis checks the stub's return against the
shape the real collaborator's return actually has at runtime.
verify_mock
Catch a lying mock. A stub that passes green but returns something the real collaborator never would. Mutation testing can't see it. The recording can.
See more →unit_test_candidates
Test what matters. Untested methods ranked by how many places call them, so the most-depended-on code gets a test first. Aim your mutation runs there too.
See more →how_to_reach
Reach any method. The collaborators and call path that get an object into the state a method needs, taken from a real run. Setup for tests that use real objects instead of doubles.
See more →split_interface
Split a fat interface. A class whose callers each use a different slice of it. It names the role interfaces to split it into, measured from who actually calls what.
See more →dead_api
Find dead public API. Public methods no caller reached in any recorded run, including the ones whose only caller is their own spec. Greyed out in your editor.
See more →
Never mock at all? Then verify_mock has nothing to check, but the fan-in ranking,
setup recipes, and interface signals still land. This complements a disciplined suite; it doesn't
replace the discipline.
No tests yet
No tests yet? Record the app instead.
Most tools need a test before they can tell you anything. Runtime Analysis records whatever runs, so point it at your app in development, staging or production and code with no specs still shows up. The more a path runs, the more we know about it, with or without a test.
On a module with zero specs you still get the real callers, the class shapes of what each method takes and returns, and a flag when a change breaks a contract the code has always honoured. When you're ready to pin it down, it scaffolds a characterization test from an actual run, so the first test isn't a guess.
$ ra observe app/billing/proration.rb # no specs in sight Proration#amount seen in 41,900 production calls callers Checkout#call · Invoice#total · Api::QuotesController returns responds to #cents, #currency (Money) raised ArgumentError on 18 calls where amount was negative # no test covers this path scaffold a characterization test from the recorded run: $ ra scaffold Proration#amount
The evidence is the traffic. A path that runs ten thousand times in production is described by those ten thousand runs, whether or not anyone wrote a test for it.
Safe to leave on in production. Under load it drops records and counts them rather than blocking your app, so the recording is honestly incomplete, never silently wrong. A bug in the tracer degrades tracing, never your app.
Plans
Free on your laptop. Hosted when you opt in. On-prem if you need it.
Recording runs on your machine and stays there. The hosted graph only sees what you choose to send it, and Enterprise puts that graph inside your own network.
Free
Every local tool, free with an account.
- The recorder, and every tool from your terminal, editor and coding agent.
- Recordings stay on your laptop. Nothing is sent to us unless you switch it on.
- Your account tells us which tools you use, never what they ran on.
Pro hosted graph
Your whole app, from every deploy.
- Opt in, and recordings from CI and production pool into one graph for your team.
- Answers from real traffic: dead code proven in production, coupling from real load, Cypher.
- Opt-in per project. Recordings hold class and method names, never your data.
Enterprise
On your terms, on your infrastructure.
- Licensing that fits your procurement and legal review.
- The hosted graph on-prem or in your own cloud, so recordings never leave your network.
- Support for rolling it out across teams.
Early access
Sign up.
Tell us what you're after and we'll email you when it's ready to try. Early sign-ups are the first group for the validation demo.
no spam · we'll email when there's something to try
We work out the interface from the messages sent to an object. What we store is class names and method names, never your data. It's inference over the call graph and your source: reliable within a method, less so once a value is passed around.
We only know the shape of what ran. A value nothing is ever called on is a blind spot, and so are values passed straight through untouched. Runtime data adds to static analysis; it doesn't replace it.
This is early. The validation demo, an agent measured with and without runtime context by false positives caught and mutants killed, is what we're building to prove it. We're not claiming it's shipped yet.