AIDive

Video pack

Claude Code verification loops: bench table, Stop hook checklist and sources

10 min read

TL;DR

  • Anthropic's own recommendation is right in spirit: the model that gets a feedback loop produces better work. The problem is who runs the loop. In 116 runs, the built-in /verify returned PASS 24 of 24 times, including on a change that broke a stated requirement.
  • A model checking its own change shares its blind spot. Haiku 4.5 missed the same unique-field case in 11 of 11 runs, in every same-model setup, and its own /verify wrote a check mark next to the wrong case.
  • On Opus 5.5, /verify cost 2.5x ($0.96 against $0.39) and changed no outcome. Plain Opus got 12 of 12 right and ran the test suite by itself in 11 of those runs.
  • A project skill is optional for the model: it fired in 10 of 24 runs, 0 of 6 on Haiku. A Stop hook fired in every run for about 20% more on Opus.
  • The check that worked came from outside the author: the same /verify run by Opus said FAIL 3 of 3 on Haiku's change and named the bug. Wired as a Stop hook, it got Haiku to fix the bug 3 of 3 times, at $1.36 per task against $0.76 for Opus writing the task alone.
  • Copy this: a Stop hook for the deterministic part, a verifier that is not the author for the judgment part, and no /verify on Opus for well-specified work.

What the measurements say

Anthropic's claim: "If Claude has that feedback loop, it will 2-3x the quality of the final result." s4 The blog turns that into a 5-step adoption workflow: pick your most repeated manual check, try the built-in /verify, write the procedure in plain English as a skill, make it deterministic, then move it to CI. s1 Since v2.1.215, "Claude no longer runs the /verify and /code-review skills on its own; invoke them with /verify or /code-review when you want them." s3

The field reports are about verification claimed but not run. Issue #96416 is a review verdict with 19 concerns, 5 verified, and an "accept as-is" issued anyway. s6 Issue #97039 is a "held up" claim after partial checks despite a loaded checklist. s7 "build passes" and "production-ready" are different bars, and each miss costs 2-3 rounds of reprompting. s8

Our bench graded every "done" with hidden acceptance tests the agent never saw. 112 task runs + 4 cross-model /verify runs = 116 runs, $66.29 of API-equivalent spend. False "done": 12 of 112, 11 of them Haiku on T6 in every setup A to E, 1 Sonnet on T6 with the skill setup, where its own new tests were failing and the skill never fired. s1

The built-in /verify said Verdict: PASS in 24 of 24 runs across the three models (23 parseable, 1 unlabeled pass), including Haiku's broken T6. On Opus it cost $0.96 against $0.39 plain (2.5x) and 125 s against 75 s; it changed no outcome. s3

The project skill written per the 5-step workflow was invoked in 10 of 24 runs: Opus 8/12, Sonnet 2/6, Haiku 0/6. The Stop hook fired in every run, at $0.47 against $0.39 on Opus (+20 %). s19

The blind spot is one requirement, T6 req 4: a single update that gives several matched documents the same unique value must raise DuplicateKeyError. Haiku missed it in 11 of 11 runs, in every setup (plain, /verify, skill, hook, tests first). Haiku's own /verify checked "update one doc onto another's email" and never tried one update hitting several documents. The tests-first rule did not help: 3/3 still miss req 4, and one also broke req 3. s1

The same Haiku change, /verify run by another model: Sonnet said PASS (1/1), Opus said FAIL 3/3, each run naming the same case, one update writing the same value to several documents, at $0.39 to $0.47 per check. As a Stop hook on Haiku (setup F), the Opus verifier got the bug fixed in 3 of 3 runs; all three hit the 60-turn cap while fixing, at an average of $1.36 per T6 run including the verifier, against $0.76 for Opus writing T6 alone with 0 false "done". s19

Measurements

Harness: Claude Code 2.1.283 headless (claude -p), isolated config dir, same prompt per task, 60-turn cap. Repo: msiemens/tinydb @ 18d73a1 (Python, 226 tests). Six feature requests T1 to T6 with 5 stated requirements each (T5: 6). Hidden acceptance tests, one per stated requirement and never shown to the agent, score every run; the repo's own suite runs too. A false "done" is a run that ended claiming completion while a hidden test or the repo suite fails.

Setups: A plain (the request only); B /verify (A, then the built-in /verify as a second turn); C verify skill (a model-invocable project skill written per the 5-step workflow); D Stop hook (prove-it.py: blocks stop while the suite is red, and blocks the first stop asking for one line of evidence per requirement); E tests first (a CLAUDE.md rule: a failing test per requirement before any code); F Opus verifier hook (claude -p /verify --model opus on the change, blocks on a non-PASS verdict, max 2 rounds).

model setup runs false done avg cost avg turns avg wall
Opus 5.5 A plain 12 0 $0.39 16.6 75 s
Opus 5.5 B /verify 12 0 $0.96 22.3 125 s
Opus 5.5 C skill 12 0 $0.43 19.5 80 s
Opus 5.5 D hook 12 0 $0.47 19.8 94 s
Opus 5.5 E tests first 2 (T6) 0 $0.64 18.5 129 s
Sonnet 5 A 7 0 $0.44 24.0 112 s
Sonnet 5 B 6 0 $1.18 36.3 210 s
Sonnet 5 C 7 1 $0.48 25.7 136 s
Sonnet 5 D 6 0 $0.53 27.0 157 s
Haiku 4.5 A 8 3 $0.34 35.4 172 s
Haiku 4.5 B 6 1 $0.68 43.3 217 s
Haiku 4.5 C 6 1 $0.26 28.2 124 s
Haiku 4.5 D 8 3 $0.37 38.9 179 s
Haiku 4.5 E 3 (T6) 3 $0.41 38.3 186 s
Haiku 4.5 F Opus verifier 3 (T6) 0 $1.36 incl. verifier 61 (cap) 457 s
Sonnet 5 F 1 (T6) 0 $2.30 incl. verifier 50 512 s

Limits: one repo (a small, well-tested Python library), six well-specified requests, 1 to 3 repetitions per cell, headless runs. The hidden tests check only what the request states.

Do this Monday

  • Write one acceptance test per stated requirement yourself, before you read the model's "done". The bench's hidden tests caught what every same-model check missed.
  • Add a Stop hook that runs your test suite and returns a block decision while it is red. It fires in every run; a skill does not.
  • Make the first stop of a task cost one line of evidence per requirement (the prove-it.py pattern): a command and its output, not a sentence.
  • Route the judgment check to a model that did not write the change: claude -p /verify --model opus on the diff, blocking on a non-PASS verdict, capped at 2 rounds.
  • On Opus 5.5 with a well-specified request, stop typing /verify by habit. It cost 2.5x and changed nothing in 12 runs; reserve it for a test-free area or a pre-existing bug hunt.
  • If you delegate to Haiku 4.5 for cost, budget the verifier: $1.36 per task with the Opus hook against $0.76 for Opus writing alone.
  • Log every retry of a failed fix in a ledger and stop the loop after a repeat, so a blocking hook cannot burn tokens on the same wrong patch.
  • Reread your requirements for the multi-row case: "one update that matches several documents" is the shape of the case 11 of 11 Haiku runs never tried.

Go further

  • The 5-step workflow and the maturity ladder, from manual check to CI gate: the top rungs (CI and PR gates) are where the deterministic part belongs once your hook works locally. s1
  • Stop hook semantics: a block decision with a reason sends the turn back to the model; read the exit code and JSON contract before writing your own gate. s19
  • Groundtruth, a Stop hook that refuses the end of the turn until checks pass: the deterministic version of the idea, ready to read as a reference implementation. s5
  • regressionledger, the cost side: a hook that stops the loop from retrying the same failed fix, the piece our Stop hook did not have when runs hit the 60-turn cap. s10

Sources

FAQ

Does the built-in /verify catch bugs?

Not in this bench. It returned PASS in 24 of 24 runs across Opus 5.5, Sonnet 5 and Haiku 4.5, including a Haiku change that broke a stated requirement. Its one useful finding was a pre-existing upstream bug unrelated to the change.

Why does a stronger verifier help a weaker writer?

Haiku's own check tested the case it had already thought of. Opus, given the same change and the same /verify, tried one update hitting several documents and said FAIL 3 of 3. Sonnet said PASS. The verifier has to see a case the author did not.

Is it cheaper to verify Haiku with Opus, or to write with Opus?

Write with Opus. The Opus verifier hook on Haiku averaged $1.36 per T6 run and hit the 60-turn cap every time; Opus writing T6 alone averaged $0.76 with 0 false "done".