TL;DR
- Anthropic's own recommendation is right in spirit: the model that gets a feedback loop produces better work. The problem is who runs the loop. In 116 runs, the built-in
/verifyreturned PASS 24 of 24 times, including on a change that broke a stated requirement. - A model checking its own change shares its blind spot. Haiku 4.5 missed the same unique-field case in 11 of 11 runs, in every same-model setup, and its own
/verifywrote a check mark next to the wrong case. - On Opus 5.5,
/verifycost 2.5x ($0.96 against $0.39) and changed no outcome. Plain Opus got 12 of 12 right and ran the test suite by itself in 11 of those runs. - A project skill is optional for the model: it fired in 10 of 24 runs, 0 of 6 on Haiku. A Stop hook fired in every run for about 20% more on Opus.
- The check that worked came from outside the author: the same
/verifyrun by Opus said FAIL 3 of 3 on Haiku's change and named the bug. Wired as a Stop hook, it got Haiku to fix the bug 3 of 3 times, at $1.36 per task against $0.76 for Opus writing the task alone. - Copy this: a Stop hook for the deterministic part, a verifier that is not the author for the judgment part, and no
/verifyon Opus for well-specified work.
What the measurements say
Anthropic's claim: "If Claude has that feedback loop, it will 2-3x the quality of the final result." s4 The blog turns that into a 5-step adoption workflow: pick your most repeated manual check, try the built-in /verify, write the procedure in plain English as a skill, make it deterministic, then move it to CI. s1 Since v2.1.215, "Claude no longer runs the /verify and /code-review skills on its own; invoke them with /verify or /code-review when you want them." s3
The field reports are about verification claimed but not run. Issue #96416 is a review verdict with 19 concerns, 5 verified, and an "accept as-is" issued anyway. s6 Issue #97039 is a "held up" claim after partial checks despite a loaded checklist. s7 "build passes" and "production-ready" are different bars, and each miss costs 2-3 rounds of reprompting. s8
Our bench graded every "done" with hidden acceptance tests the agent never saw. 112 task runs + 4 cross-model /verify runs = 116 runs, $66.29 of API-equivalent spend. False "done": 12 of 112, 11 of them Haiku on T6 in every setup A to E, 1 Sonnet on T6 with the skill setup, where its own new tests were failing and the skill never fired. s1
The built-in /verify said Verdict: PASS in 24 of 24 runs across the three models (23 parseable, 1 unlabeled pass), including Haiku's broken T6. On Opus it cost $0.96 against $0.39 plain (2.5x) and 125 s against 75 s; it changed no outcome. s3
The project skill written per the 5-step workflow was invoked in 10 of 24 runs: Opus 8/12, Sonnet 2/6, Haiku 0/6. The Stop hook fired in every run, at $0.47 against $0.39 on Opus (+20 %). s19
The blind spot is one requirement, T6 req 4: a single update that gives several matched documents the same unique value must raise DuplicateKeyError. Haiku missed it in 11 of 11 runs, in every setup (plain, /verify, skill, hook, tests first). Haiku's own /verify checked "update one doc onto another's email" and never tried one update hitting several documents. The tests-first rule did not help: 3/3 still miss req 4, and one also broke req 3. s1
The same Haiku change, /verify run by another model: Sonnet said PASS (1/1), Opus said FAIL 3/3, each run naming the same case, one update writing the same value to several documents, at $0.39 to $0.47 per check. As a Stop hook on Haiku (setup F), the Opus verifier got the bug fixed in 3 of 3 runs; all three hit the 60-turn cap while fixing, at an average of $1.36 per T6 run including the verifier, against $0.76 for Opus writing T6 alone with 0 false "done". s19
Measurements
Harness: Claude Code 2.1.283 headless (claude -p), isolated config dir, same prompt per task, 60-turn cap. Repo: msiemens/tinydb @ 18d73a1 (Python, 226 tests). Six feature requests T1 to T6 with 5 stated requirements each (T5: 6). Hidden acceptance tests, one per stated requirement and never shown to the agent, score every run; the repo's own suite runs too. A false "done" is a run that ended claiming completion while a hidden test or the repo suite fails.
Setups: A plain (the request only); B /verify (A, then the built-in /verify as a second turn); C verify skill (a model-invocable project skill written per the 5-step workflow); D Stop hook (prove-it.py: blocks stop while the suite is red, and blocks the first stop asking for one line of evidence per requirement); E tests first (a CLAUDE.md rule: a failing test per requirement before any code); F Opus verifier hook (claude -p /verify --model opus on the change, blocks on a non-PASS verdict, max 2 rounds).
| model | setup | runs | false done | avg cost | avg turns | avg wall |
|---|---|---|---|---|---|---|
| Opus 5.5 | A plain | 12 | 0 | $0.39 | 16.6 | 75 s |
| Opus 5.5 | B /verify | 12 | 0 | $0.96 | 22.3 | 125 s |
| Opus 5.5 | C skill | 12 | 0 | $0.43 | 19.5 | 80 s |
| Opus 5.5 | D hook | 12 | 0 | $0.47 | 19.8 | 94 s |
| Opus 5.5 | E tests first | 2 (T6) | 0 | $0.64 | 18.5 | 129 s |
| Sonnet 5 | A | 7 | 0 | $0.44 | 24.0 | 112 s |
| Sonnet 5 | B | 6 | 0 | $1.18 | 36.3 | 210 s |
| Sonnet 5 | C | 7 | 1 | $0.48 | 25.7 | 136 s |
| Sonnet 5 | D | 6 | 0 | $0.53 | 27.0 | 157 s |
| Haiku 4.5 | A | 8 | 3 | $0.34 | 35.4 | 172 s |
| Haiku 4.5 | B | 6 | 1 | $0.68 | 43.3 | 217 s |
| Haiku 4.5 | C | 6 | 1 | $0.26 | 28.2 | 124 s |
| Haiku 4.5 | D | 8 | 3 | $0.37 | 38.9 | 179 s |
| Haiku 4.5 | E | 3 (T6) | 3 | $0.41 | 38.3 | 186 s |
| Haiku 4.5 | F Opus verifier | 3 (T6) | 0 | $1.36 incl. verifier | 61 (cap) | 457 s |
| Sonnet 5 | F | 1 (T6) | 0 | $2.30 incl. verifier | 50 | 512 s |
Limits: one repo (a small, well-tested Python library), six well-specified requests, 1 to 3 repetitions per cell, headless runs. The hidden tests check only what the request states.
Do this Monday
- Write one acceptance test per stated requirement yourself, before you read the model's "done". The bench's hidden tests caught what every same-model check missed.
- Add a Stop hook that runs your test suite and returns a block decision while it is red. It fires in every run; a skill does not.
- Make the first stop of a task cost one line of evidence per requirement (the
prove-it.pypattern): a command and its output, not a sentence. - Route the judgment check to a model that did not write the change:
claude -p /verify --model opuson the diff, blocking on a non-PASS verdict, capped at 2 rounds. - On Opus 5.5 with a well-specified request, stop typing
/verifyby habit. It cost 2.5x and changed nothing in 12 runs; reserve it for a test-free area or a pre-existing bug hunt. - If you delegate to Haiku 4.5 for cost, budget the verifier: $1.36 per task with the Opus hook against $0.76 for Opus writing alone.
- Log every retry of a failed fix in a ledger and stop the loop after a repeat, so a blocking hook cannot burn tokens on the same wrong patch.
- Reread your requirements for the multi-row case: "one update that matches several documents" is the shape of the case 11 of 11 Haiku runs never tried.
Go further
- The 5-step workflow and the maturity ladder, from manual check to CI gate: the top rungs (CI and PR gates) are where the deterministic part belongs once your hook works locally. s1
- Stop hook semantics: a block decision with a reason sends the turn back to the model; read the exit code and JSON contract before writing your own gate. s19
- Groundtruth, a Stop hook that refuses the end of the turn until checks pass: the deterministic version of the idea, ready to read as a reference implementation. s5
- regressionledger, the cost side: a hook that stops the loop from retrying the same failed fix, the piece our Stop hook did not have when runs hit the 60-turn cap. s10
Sources
- Building verification loops in Claude Code with skills, Anthropic blog. Why read it: the 5-step workflow and the ladder the bench's skill setup was built on.
- Building verification loops in Claude Code, official Claude channel. Why read it: three minutes on what
/verifydoes on first run, before you decide whether to keep it. - Claude Code CHANGELOG, GitHub. Why read it: v2.1.215 is where
/verifybecame user-invoked only, which changes how often it runs for you. - Boris Cherny: give Claude a way to verify its work, X. Why read it: the "2-3x the quality" claim, in its exact wording.
- Groundtruth, GitHub. Why read it: a working Stop hook gate to copy from instead of writing yours from zero.
- Issue #96416, GitHub. Why read it: a dated transcript of a review verdict that verified 5 of 19 concerns and accepted anyway.
- Issue #97039, GitHub. Why read it: the same failure two days later with a checklist loaded, so the checklist is not the fix.
- AI coding agents can verify some of their work now, dev.to. Why read it: the clearest statement of the gap between "build passes" and "production-ready".
- Saguaro, GitHub. Why read it: the in-loop versus PR-level review debate, with the counterargument in the comments.
- regressionledger, GitHub. Why read it: the retry-cost problem a blocking hook creates, and one way to cap it.
- SPICE simulation to oscilloscope to verification with Claude Code, personal blog. Why read it: a verification loop whose oracle is a physical instrument.
- Hooks reference, code.claude.com. Why read it: the block decision contract your Stop hook has to honour.
FAQ
Does the built-in /verify catch bugs?
Not in this bench. It returned PASS in 24 of 24 runs across Opus 5.5, Sonnet 5 and Haiku 4.5, including a Haiku change that broke a stated requirement. Its one useful finding was a pre-existing upstream bug unrelated to the change.
Why does a stronger verifier help a weaker writer?
Haiku's own check tested the case it had already thought of. Opus, given the same change and the same /verify, tried one update hitting several documents and said FAIL 3 of 3. Sonnet said PASS. The verifier has to see a case the author did not.
Is it cheaper to verify Haiku with Opus, or to write with Opus?
Write with Opus. The Opus verifier hook on Haiku averaged $1.36 per T6 run and hit the 60-turn cap every time; Opus writing T6 alone averaged $0.76 with 0 false "done".
AIDive