AIDive

Stop Letting Claude Code Grade Its Own Work

By AIDive · Published

Coding agentsAI models

It graded itself: PASS

Claude Code's built-in verification check, /verify, returned PASS on a feature that was broken. That is the result of a bench of 116 Claude Code sessions on a real repository, where every "done" was graded afterwards by acceptance tests the agent never saw.

Boris Cherny, who created Claude Code, calls verification the most important thing you can give it. The built-in check has only run on demand since Claude Code v2.1.215: you type /verify, or it does not run. On this bench it said PASS in 24 runs out of 24, and one of those runs had shipped a broken feature. What actually caught the bug is covered below, along with the setup that cost two and a half times more and changed nothing.

The bench: 116 runs, tests it never saw

A false "done" is a session that ends by claiming the work is finished while a hidden acceptance test, or the repository's own suite, fails. The bench measures how often that happens under six verification setups.

The doubt behind it comes from Claude Code's own issue tracker: issue #96416, filed on 2026-09-23, describes a review that listed 19 concerns, verified 5 of them and still concluded "accept as is".

Parameter Value
Repository msiemens/tinydb (Python document database), commit 18d73a1
Repository test suite 226 tests
Feature requests 6, each stating 5 requirements
Hidden acceptance tests one per stated requirement, written before any run, never shown to the agent
Models Opus 5.5, Sonnet 5, Haiku 4.5
Claude Code 2.1.283, headless (claude -p), 60-turn cap
Sessions 112 task runs + 4 cross-model /verify runs = 116
Spend $66.29 API-equivalent

The six setups run from plain Claude Code (the request only) to a second model checking the work at the stop:

Setup What it adds
A plain nothing
B /verify the built-in /verify typed as a second turn
C verify skill a project skill written per Anthropic's workflow
D Stop hook a script that blocks the stop while the suite is red and asks for evidence per requirement
E tests first a CLAUDE.md rule: a failing test per requirement before any code
F Opus verifier a Stop hook that runs /verify with Opus on the change

Clear requests on a well-tested library are the easy case. Plain Opus 5.5 got all 12 of its runs right, and it ran the test suite by itself before saying done in 11 of them.

The built-in /verify: PASS, every time, 2.5x

/verify is the verification skill that ships with Claude Code. It runs the change, reads the diff and writes a step-by-step verdict. Since v2.1.215 it is user-invoked only, so on the bench it was typed after every task as a second turn.

It returned a PASS verdict in 24 runs out of 24 across the three models, including Haiku's broken change. On Opus 5.5 it changed no outcome and cost 2.5 times as much:

Opus 5.5, per task Plain With /verify
Cost $0.39 $0.96
Wall time 75 s 125 s
Outcomes changed 0 of 12

To be fair, /verify found one real bug: on one Opus task it flagged a reuse-after-close problem that was already in the library, not caused by the change it was checking. On work that was already right, it is an expensive second opinion.

Skill or hook: the one that gets skipped

A skill is a markdown procedure Claude can open when it judges it relevant. Anthropic's blog post on verification loops describes a six-step recipe: pick the manual follow-up you do most often, try the built-in /verify first, write the procedure in plain English, turn it into a skill, then invoke it on a new task and iterate. The bench's skill, verify-change, was built exactly that way and tells Claude to prove every requirement against the real code before saying done.

A skill is a suggestion, and Claude decides whether to open it:

Model Sessions that opened the skill
Opus 5.5 8 of 12
Sonnet 5 2 of 6
Haiku 4.5 0 of 6
Total 10 of 24

The one Sonnet false "done" came from a session where the skill never opened: it ended with two of its own new tests failing.

A Stop hook is a script Claude Code runs every time the agent tries to finish. It can refuse the stop and send the agent back with a reason. The bench's hook runs the test suite, blocks while it is red, and on the first stop asks for one line of evidence per requirement. It fired in every session. On Opus it cost 20% more, $0.47 against $0.39 per task (94 s against 75 s). In this bench it never met a failing suite at stop time, so it had nothing to catch: a hook always fires, but it only checks what you told it to check.

The blind spot: 11 of 11 missed

The failing task asked for unique fields on a table: two users cannot share an email, and any update that would leave two documents with the same unique value must raise DuplicateKeyError.

Haiku 4.5 tested moving one user onto another user's email, and that was refused correctly. It never tested one update that matches several documents and writes the same new email to all of them. That case went through without an error in 11 Haiku sessions out of 11, in every same-model setup: plain, /verify, the skill, the Stop hook and tests first.

Haiku's own /verify ticked the case it had tried and wrote PASS. The hidden test reported DID NOT RAISE DuplicateKeyError. Sonnet 5, asked to run /verify on the same change, also returned PASS. Opus and Sonnet both wrote this feature correctly on their own, so this is one model on one task, but a check written by the same model shares its blind spot.

The outside check: Opus says FAIL

The same /verify on the same Haiku change, run by Opus 5.5, returned FAIL in 3 runs out of 3, each time naming the missed case: one update that matches several documents writes the same value to all of them with no error. The catch came from a different model, not from the author and not from the author's own check.

Wired as a Stop hook (setup F), Opus checks Haiku's work every time Haiku tries to finish. The results:

T6, per run Haiku + Opus checker Opus 5.5 alone
Bug fixed / false "done" fixed in 3 of 3 0 false "done"
Turns 60-turn cap hit in every run
Cost $1.36 including the checker $0.76

The outside check works. On this task it cost more than the stronger model writing the feature alone.

What to copy, and what it costs

Stop letting the model that wrote a change be the only one that checks it.

Rule Why Cost on this bench
Keep a Stop hook for anything a script can check it fires in every session about 20% more on Opus
Make the real check come from outside the author: a stronger model at the gate, or your own tests written from the request the same model missed its own bug 11 times out of 11 $1.36 per task for Haiku + Opus checker
On a clear request with Opus, skip typing /verify 0 outcomes changed 2.5x cost, 125 s against 75 s

The limits: one small library, six clear requests, one to three runs per cell, headless sessions, and hidden tests that check only what each request states. On messier work, the numbers will move.

Sources

Frequently asked questions

Does Claude Code's /verify catch bugs?
Not reliably when the same model checks its own work. On a 116-session bench, /verify returned PASS in 24 of 24 runs, including a Haiku 4.5 change that failed a stated requirement; run by Opus 5.5 on that same change, it returned FAIL 3 of 3.
Is /verify worth it on Claude Code with Opus?
Not on clear requests. On Opus 5.5 it raised the cost per task from $0.39 to $0.96 (2.5x) and the time from 75 s to 125 s, and it changed none of the 12 outcomes, because plain Opus already ran the test suite itself in 11 of 12 runs.
Should I use a Claude Code skill or a Stop hook for verification?
A Stop hook, for anything a script can check. Claude opened a verification skill in only 10 of 24 sessions (Opus 8 of 12, Sonnet 2 of 6, Haiku 0 of 6), while a Stop hook runs every time the agent tries to finish; it cost about 20% more on Opus.
What is a Claude Code Stop hook?
A script Claude Code runs every time the agent tries to finish. It can refuse the stop with a JSON decision of block and a reason, which sends the agent back to work; it checks only what the script tests.
Can a stronger model check a weaker model's code in Claude Code?
Yes. An Opus 5.5 /verify wired as a Stop hook made Haiku 4.5 fix its missed case in 3 of 3 runs. It was costly: every run hit the 60-turn cap and averaged $1.36, against $0.76 for Opus writing the feature alone.
Why does an AI model miss bugs in its own code?
Its check tests the cases it already thought of. Haiku verified moving one user onto another's email but never one update hitting several documents, so its own /verify ticked the tried case and wrote PASS while the hidden test failed.

Related videos