# Stop Letting Claude Code Grade Its Own Work

Video: https://www.youtube.com/watch?v=oL2gYslpkJw
Article: https://aidive.dev/videos/claude-code-verification-loops/
Published: 2026-09-29

## Chapters

- [0:00](https://www.youtube.com/watch?v=oL2gYslpkJw&t=0s) It graded itself: PASS
- [0:32](https://www.youtube.com/watch?v=oL2gYslpkJw&t=32s) The bench: 116 runs, tests it never saw
- [1:35](https://www.youtube.com/watch?v=oL2gYslpkJw&t=95s) The built-in /verify: PASS, every time, 2.5x
- [2:16](https://www.youtube.com/watch?v=oL2gYslpkJw&t=136s) Skill or hook: the one that gets skipped
- [3:04](https://www.youtube.com/watch?v=oL2gYslpkJw&t=184s) The blind spot: 11 of 11 missed
- [3:49](https://www.youtube.com/watch?v=oL2gYslpkJw&t=229s) The outside check: Opus says FAIL
- [4:29](https://www.youtube.com/watch?v=oL2gYslpkJw&t=269s) What to copy, and what it costs

## TL;DR

- Claude Code's built-in /verify returned PASS in 24 of 24 runs on a 116-session bench, including a Haiku 4.5 change that broke a stated requirement.
- The model that wrote a change shares its blind spot: Haiku missed one unique-field case in 11 of 11 runs in every same-model setup, and /verify run by Haiku or Sonnet said PASS, while Opus 5.5 said FAIL 3 of 3 and named the bug.
- On Opus 5.5, typing /verify after each task cost 2.5x ($0.96 against $0.39) and 125 s against 75 s, and changed no outcome; plain Opus got 12 of 12 right and ran the tests itself in 11.
- A project verification skill opened in only 10 of 24 sessions (Haiku never); a Stop hook fired in every session for about 20% more on Opus.
- An Opus checker wired as a Stop hook made Haiku fix the bug 3 of 3 times, but every run hit the 60-turn cap and cost $1.36, against $0.76 for Opus writing the task alone.

## It graded itself: PASS

Claude Code's built-in verification check, `/verify`, returned PASS on a feature that was broken. That is the result of a bench of 116 Claude Code sessions on a real repository, where every "done" was graded afterwards by acceptance tests the agent never saw.

Boris Cherny, who created Claude Code, calls verification the most important thing you can give it. The built-in check has only run on demand since Claude Code v2.1.215: you type `/verify`, or it does not run. On this bench it said PASS in 24 runs out of 24, and one of those runs had shipped a broken feature. What actually caught the bug is covered below, along with the setup that cost two and a half times more and changed nothing.

## The bench: 116 runs, tests it never saw

A false "done" is a session that ends by claiming the work is finished while a hidden acceptance test, or the repository's own suite, fails. The bench measures how often that happens under six verification setups.

The doubt behind it comes from Claude Code's own issue tracker: issue #96416, filed on 2026-09-23, describes a review that listed 19 concerns, verified 5 of them and still concluded "accept as is".

| Parameter | Value |
|---|---|
| Repository | msiemens/tinydb (Python document database), commit 18d73a1 |
| Repository test suite | 226 tests |
| Feature requests | 6, each stating 5 requirements |
| Hidden acceptance tests | one per stated requirement, written before any run, never shown to the agent |
| Models | Opus 5.5, Sonnet 5, Haiku 4.5 |
| Claude Code | 2.1.283, headless (`claude -p`), 60-turn cap |
| Sessions | 112 task runs + 4 cross-model `/verify` runs = 116 |
| Spend | $66.29 API-equivalent |

The six setups run from plain Claude Code (the request only) to a second model checking the work at the stop:

| Setup | What it adds |
|---|---|
| A plain | nothing |
| B /verify | the built-in `/verify` typed as a second turn |
| C verify skill | a project skill written per Anthropic's workflow |
| D Stop hook | a script that blocks the stop while the suite is red and asks for evidence per requirement |
| E tests first | a CLAUDE.md rule: a failing test per requirement before any code |
| F Opus verifier | a Stop hook that runs `/verify` with Opus on the change |

Clear requests on a well-tested library are the easy case. Plain Opus 5.5 got all 12 of its runs right, and it ran the test suite by itself before saying done in 11 of them.

## The built-in /verify: PASS, every time, 2.5x

`/verify` is the verification skill that ships with Claude Code. It runs the change, reads the diff and writes a step-by-step verdict. Since v2.1.215 it is user-invoked only, so on the bench it was typed after every task as a second turn.

It returned a PASS verdict in 24 runs out of 24 across the three models, including Haiku's broken change. On Opus 5.5 it changed no outcome and cost 2.5 times as much:

| Opus 5.5, per task | Plain | With /verify |
|---|---|---|
| Cost | $0.39 | $0.96 |
| Wall time | 75 s | 125 s |
| Outcomes changed | | 0 of 12 |

To be fair, `/verify` found one real bug: on one Opus task it flagged a reuse-after-close problem that was already in the library, not caused by the change it was checking. On work that was already right, it is an expensive second opinion.

## Skill or hook: the one that gets skipped

A skill is a markdown procedure Claude can open when it judges it relevant. Anthropic's blog post on verification loops describes a six-step recipe: pick the manual follow-up you do most often, try the built-in `/verify` first, write the procedure in plain English, turn it into a skill, then invoke it on a new task and iterate. The bench's skill, `verify-change`, was built exactly that way and tells Claude to prove every requirement against the real code before saying done.

A skill is a suggestion, and Claude decides whether to open it:

| Model | Sessions that opened the skill |
|---|---|
| Opus 5.5 | 8 of 12 |
| Sonnet 5 | 2 of 6 |
| Haiku 4.5 | 0 of 6 |
| Total | 10 of 24 |

The one Sonnet false "done" came from a session where the skill never opened: it ended with two of its own new tests failing.

A Stop hook is a script Claude Code runs every time the agent tries to finish. It can refuse the stop and send the agent back with a reason. The bench's hook runs the test suite, blocks while it is red, and on the first stop asks for one line of evidence per requirement. It fired in every session. On Opus it cost 20% more, $0.47 against $0.39 per task (94 s against 75 s). In this bench it never met a failing suite at stop time, so it had nothing to catch: a hook always fires, but it only checks what you told it to check.

## The blind spot: 11 of 11 missed

The failing task asked for unique fields on a table: two users cannot share an email, and any update that would leave two documents with the same unique value must raise `DuplicateKeyError`.

Haiku 4.5 tested moving one user onto another user's email, and that was refused correctly. It never tested one update that matches several documents and writes the same new email to all of them. That case went through without an error in 11 Haiku sessions out of 11, in every same-model setup: plain, `/verify`, the skill, the Stop hook and tests first.

Haiku's own `/verify` ticked the case it had tried and wrote PASS. The hidden test reported `DID NOT RAISE DuplicateKeyError`. Sonnet 5, asked to run `/verify` on the same change, also returned PASS. Opus and Sonnet both wrote this feature correctly on their own, so this is one model on one task, but a check written by the same model shares its blind spot.

## The outside check: Opus says FAIL

The same `/verify` on the same Haiku change, run by Opus 5.5, returned FAIL in 3 runs out of 3, each time naming the missed case: one update that matches several documents writes the same value to all of them with no error. The catch came from a different model, not from the author and not from the author's own check.

Wired as a Stop hook (setup F), Opus checks Haiku's work every time Haiku tries to finish. The results:

| T6, per run | Haiku + Opus checker | Opus 5.5 alone |
|---|---|---|
| Bug fixed / false "done" | fixed in 3 of 3 | 0 false "done" |
| Turns | 60-turn cap hit in every run | |
| Cost | $1.36 including the checker | $0.76 |

The outside check works. On this task it cost more than the stronger model writing the feature alone.

## What to copy, and what it costs

Stop letting the model that wrote a change be the only one that checks it.

| Rule | Why | Cost on this bench |
|---|---|---|
| Keep a Stop hook for anything a script can check | it fires in every session | about 20% more on Opus |
| Make the real check come from outside the author: a stronger model at the gate, or your own tests written from the request | the same model missed its own bug 11 times out of 11 | $1.36 per task for Haiku + Opus checker |
| On a clear request with Opus, skip typing `/verify` | 0 outcomes changed | 2.5x cost, 125 s against 75 s |

The limits: one small library, six clear requests, one to three runs per cell, headless sessions, and hidden tests that check only what each request states. On messier work, the numbers will move.

## FAQ

### Does Claude Code's /verify catch bugs?

Not reliably when the same model checks its own work. On a 116-session bench, /verify returned PASS in 24 of 24 runs, including a Haiku 4.5 change that failed a stated requirement; run by Opus 5.5 on that same change, it returned FAIL 3 of 3.

### Is /verify worth it on Claude Code with Opus?

Not on clear requests. On Opus 5.5 it raised the cost per task from $0.39 to $0.96 (2.5x) and the time from 75 s to 125 s, and it changed none of the 12 outcomes, because plain Opus already ran the test suite itself in 11 of 12 runs.

### Should I use a Claude Code skill or a Stop hook for verification?

A Stop hook, for anything a script can check. Claude opened a verification skill in only 10 of 24 sessions (Opus 8 of 12, Sonnet 2 of 6, Haiku 0 of 6), while a Stop hook runs every time the agent tries to finish; it cost about 20% more on Opus.

### What is a Claude Code Stop hook?

A script Claude Code runs every time the agent tries to finish. It can refuse the stop with a JSON decision of block and a reason, which sends the agent back to work; it checks only what the script tests.

### Can a stronger model check a weaker model's code in Claude Code?

Yes. An Opus 5.5 /verify wired as a Stop hook made Haiku 4.5 fix its missed case in 3 of 3 runs. It was costly: every run hit the 60-turn cap and averaged $1.36, against $0.76 for Opus writing the feature alone.

### Why does an AI model miss bugs in its own code?

Its check tests the cases it already thought of. Haiku verified moving one user onto another's email but never one update hitting several documents, so its own /verify ticked the tried case and wrote PASS while the hidden test failed.

## Sources

- [Building verification loops in Claude Code with skills](https://claude.com/blog/building-verification-loops-in-claude-code-with-skills) — Anthropic
- [Hooks reference](https://code.claude.com/docs/en/hooks) — Anthropic
- [Claude Code CHANGELOG](https://github.com/anthropics/claude-code/blob/main/CHANGELOG.md) — Anthropic
- [Issue #96416: Review verdicts assert verification that was never performed](https://github.com/anthropics/claude-code/issues/96416) — GitHub
- [TinyDB](https://github.com/msiemens/tinydb) — GitHub
- [Boris Cherny on verification](https://x.com/bcherny/status/2007179861115511237) — X
