It graded itself: PASS
Claude Code's built-in verification check, /verify, returned PASS on a feature that was broken. That is the result of a bench of 116 Claude Code sessions on a real repository, where every "done" was graded afterwards by acceptance tests the agent never saw.
Boris Cherny, who created Claude Code, calls verification the most important thing you can give it. The built-in check has only run on demand since Claude Code v2.1.215: you type /verify, or it does not run. On this bench it said PASS in 24 runs out of 24, and one of those runs had shipped a broken feature. What actually caught the bug is covered below, along with the setup that cost two and a half times more and changed nothing.
The bench: 116 runs, tests it never saw
A false "done" is a session that ends by claiming the work is finished while a hidden acceptance test, or the repository's own suite, fails. The bench measures how often that happens under six verification setups.
The doubt behind it comes from Claude Code's own issue tracker: issue #96416, filed on 2026-09-23, describes a review that listed 19 concerns, verified 5 of them and still concluded "accept as is".
| Parameter | Value |
|---|---|
| Repository | msiemens/tinydb (Python document database), commit 18d73a1 |
| Repository test suite | 226 tests |
| Feature requests | 6, each stating 5 requirements |
| Hidden acceptance tests | one per stated requirement, written before any run, never shown to the agent |
| Models | Opus 5.5, Sonnet 5, Haiku 4.5 |
| Claude Code | 2.1.283, headless (claude -p), 60-turn cap |
| Sessions | 112 task runs + 4 cross-model /verify runs = 116 |
| Spend | $66.29 API-equivalent |
The six setups run from plain Claude Code (the request only) to a second model checking the work at the stop:
| Setup | What it adds |
|---|---|
| A plain | nothing |
| B /verify | the built-in /verify typed as a second turn |
| C verify skill | a project skill written per Anthropic's workflow |
| D Stop hook | a script that blocks the stop while the suite is red and asks for evidence per requirement |
| E tests first | a CLAUDE.md rule: a failing test per requirement before any code |
| F Opus verifier | a Stop hook that runs /verify with Opus on the change |
Clear requests on a well-tested library are the easy case. Plain Opus 5.5 got all 12 of its runs right, and it ran the test suite by itself before saying done in 11 of them.
The built-in /verify: PASS, every time, 2.5x
/verify is the verification skill that ships with Claude Code. It runs the change, reads the diff and writes a step-by-step verdict. Since v2.1.215 it is user-invoked only, so on the bench it was typed after every task as a second turn.
It returned a PASS verdict in 24 runs out of 24 across the three models, including Haiku's broken change. On Opus 5.5 it changed no outcome and cost 2.5 times as much:
| Opus 5.5, per task | Plain | With /verify |
|---|---|---|
| Cost | $0.39 | $0.96 |
| Wall time | 75 s | 125 s |
| Outcomes changed | 0 of 12 |
To be fair, /verify found one real bug: on one Opus task it flagged a reuse-after-close problem that was already in the library, not caused by the change it was checking. On work that was already right, it is an expensive second opinion.
Skill or hook: the one that gets skipped
A skill is a markdown procedure Claude can open when it judges it relevant. Anthropic's blog post on verification loops describes a six-step recipe: pick the manual follow-up you do most often, try the built-in /verify first, write the procedure in plain English, turn it into a skill, then invoke it on a new task and iterate. The bench's skill, verify-change, was built exactly that way and tells Claude to prove every requirement against the real code before saying done.
A skill is a suggestion, and Claude decides whether to open it:
| Model | Sessions that opened the skill |
|---|---|
| Opus 5.5 | 8 of 12 |
| Sonnet 5 | 2 of 6 |
| Haiku 4.5 | 0 of 6 |
| Total | 10 of 24 |
The one Sonnet false "done" came from a session where the skill never opened: it ended with two of its own new tests failing.
A Stop hook is a script Claude Code runs every time the agent tries to finish. It can refuse the stop and send the agent back with a reason. The bench's hook runs the test suite, blocks while it is red, and on the first stop asks for one line of evidence per requirement. It fired in every session. On Opus it cost 20% more, $0.47 against $0.39 per task (94 s against 75 s). In this bench it never met a failing suite at stop time, so it had nothing to catch: a hook always fires, but it only checks what you told it to check.
The blind spot: 11 of 11 missed
The failing task asked for unique fields on a table: two users cannot share an email, and any update that would leave two documents with the same unique value must raise DuplicateKeyError.
Haiku 4.5 tested moving one user onto another user's email, and that was refused correctly. It never tested one update that matches several documents and writes the same new email to all of them. That case went through without an error in 11 Haiku sessions out of 11, in every same-model setup: plain, /verify, the skill, the Stop hook and tests first.
Haiku's own /verify ticked the case it had tried and wrote PASS. The hidden test reported DID NOT RAISE DuplicateKeyError. Sonnet 5, asked to run /verify on the same change, also returned PASS. Opus and Sonnet both wrote this feature correctly on their own, so this is one model on one task, but a check written by the same model shares its blind spot.
The outside check: Opus says FAIL
The same /verify on the same Haiku change, run by Opus 5.5, returned FAIL in 3 runs out of 3, each time naming the missed case: one update that matches several documents writes the same value to all of them with no error. The catch came from a different model, not from the author and not from the author's own check.
Wired as a Stop hook (setup F), Opus checks Haiku's work every time Haiku tries to finish. The results:
| T6, per run | Haiku + Opus checker | Opus 5.5 alone |
|---|---|---|
| Bug fixed / false "done" | fixed in 3 of 3 | 0 false "done" |
| Turns | 60-turn cap hit in every run | |
| Cost | $1.36 including the checker | $0.76 |
The outside check works. On this task it cost more than the stronger model writing the feature alone.
What to copy, and what it costs
Stop letting the model that wrote a change be the only one that checks it.
| Rule | Why | Cost on this bench |
|---|---|---|
| Keep a Stop hook for anything a script can check | it fires in every session | about 20% more on Opus |
| Make the real check come from outside the author: a stronger model at the gate, or your own tests written from the request | the same model missed its own bug 11 times out of 11 | $1.36 per task for Haiku + Opus checker |
On a clear request with Opus, skip typing /verify |
0 outcomes changed | 2.5x cost, 125 s against 75 s |
The limits: one small library, six clear requests, one to three runs per cell, headless sessions, and hidden tests that check only what each request states. On messier work, the numbers will move.
AIDive