AIDive

Anthropic Wrote A Claude Code Playbook. Nobody Measured It

By AIDive · Published

Coding agentsAutomation & workflows

The playbook nobody measured

Anthropic published an AI-native SDLC playbook for Claude Code: six stages, each ending in a committed file, taught as a free course. Its central claim is that code is no longer the bottleneck, and that the chain of committed artifacts manages what is. The document itself carries no measurement of any kind — no times, no costs, no benchmark. We ran the first timed test on a real repository: sixteen timed Claude Code sessions, every gate priced, and a verdict that splits the chain down the middle. Along the way, a two-minute fix pushed through the full chain priced the ceremony, and our own deploy was stopped twice, once by four lines of shell.

The playbook and the rig

The test rig: sixteen timed Claude Code sessions, about $12 of compute, one real repository. The playbook runs six stages — plan, design, build, test, deploy, maintain. Every stage ends in a committed file, and the stage after it reads that file: intent, spec, plan, the pull request, the incident record. The commits are the audit trail. Anthropic teaches it as a free 14-lesson course, about an hour, written for enterprises with review gates; we tested what survives contact with one developer.

The repository is the RealWorld demo app (Express, TypeScript, Prisma, Postgres) — a real project with real tests, and one suite broken on a fresh clone (four passing suites, 14 green tests, two seconds to run). That bug becomes the control group later. The scoring rule: a gate pays when its output changes what ships, for less than it costs.

The entry ticket is a memory file at the repo root — commands, conventions, architecture, the mistakes the model keeps repeating, kept under a page. Ours was written and committed in 63 seconds for $0.44. One honest caveat on method: headless runs compress the playbook's interviews into single prompts.

Plan: intent.md in twenty-nine seconds

The first gate captures the idea before anyone designs anything. The feature request: readers want to mute authors who flood their feed. The playbook calls the result a proto-spec — written with the model, owned by you — and allows three origins: an idea, a filed ticket, or an incident alert. The template has five sections whose headings do the thinking: problem, proposed outcome, affected users and systems, constraints, open questions. The working loop is five moves: describe, brainstorm, generate from the template, correct, commit.

intent.md time cost
Mute-authors feature 29 s $0.18
Broken test suite 39 s —

The value sits at the bottom, in the open questions: what happens to favorites from a muted author? Do their pages stay reachable? These are decisions a coding agent would otherwise make silently, now written down and dated. The file is committed, so authorship and timestamp survive the chat, and the product owner corrects the draft before accepting it. Anthropic's own target for this stage is elicitation in hours, not weeks; solo, it is under a minute.

Design: the spec flags its own prerequisite

Gate two turns the intent into a spec with one prompt from the course: read the intent, produce a requirements and design spec, apply the available skills — the skills that are supposed to carry your brand, security and UX policy. Two minutes later we had roughly 2,300 words of competent spec: endpoints, data model, feed behavior, edge cases. It even recorded what it could not satisfy, exactly as the prompt asks.

The twist sits in its flagged concerns, written by the model itself: "C0. No org skills available. This spec has not been checked against any policy." The stage's whole premise assumed files that do not exist in most setups — the explainer videos skip that prerequisite; the agent put it in writing. The second flag was tamer: open-question defaults need product sign-off before build.

The lesson is strict about pairing (spec and intent commit together, a human approves the move to build), and there is a reading bill: about 12 minutes of product-owner time per spec. The playbook even tracks rework — spec commits dated after build starts count against you. On a team that encoded its policies, this gate is where they execute. Solo, you are paying for a promise the setup cannot keep yet.

Build: plan mode, TDD, and what the loop really checks

Gate three is Plan Mode, and the bar is brutal and useful: an engineer who never saw the conversation could implement from the plan alone. Plan Mode enforces the reading half itself — the model cannot edit files until the plan is accepted. Ours came out at about 4,000 words in four minutes, naming the files that change, the order of work, the risks and the proof, and recording three labelled deviations from the spec that come back later in review.

The build runs on a loop: write the failing test, make it pass, one target, all green or the task is not done. The loop is protected (an agent fixing code must not weaken the check on that code) and paired with a verifier — a second check in a fresh context, unswayed by the session that wrote the code.

Build result value
Agent time ~9 min, 91 turns
Cost ~$2
Change 15 files, Mutes table, two endpoints, both feeds filtered
Tests 5 suites, 50 tests, all green on an independent rerun
First-pass merge yes

The asterisk: green proves what the loop contains, nothing more. End-to-end was never run — it needs a live server and a seeded database, and a loop aimed at stale fakes would glow just as green. Team scale adds parallel sessions in worktrees (two or three is the stated ceiling); we did not test it.

Deploy: the review, and the gate that said no

The deploy gate has two layers, and both said no to us. Layer one reads the diff under a written policy at the repo root: three passes (bugs, security, compliance against the spec and plan), "Important" reserved for broken behavior, leaked data or a breached policy, five nits maximum with the rest summarized as a count — the policy caps its own noise. Two minutes of review, $0.80, and it ran the real checks: tests, build, lint against the plan's recorded baseline, formatting across nine files. Verdict: zero Important findings, six nits, one over the cap summarized. It closed with a line we did not prompt: this agent does not approve — approval stays with a human code owner behind branch protection.

Layer two is the gate itself. We asked for the deploy; the model refused on its own, because the feature was not on the shipping branch. That is judgment, not enforcement. So we merged and asked again: four lines of shell answered in 14 seconds — blocked, release authorization required. Exit code 2 stops the tool call and the reason lands back with the model. The pipeline side got the same treatment: a broken build triaged headless in 11 seconds for $0.13 — it read the log, named the exact cause, and proposed the diff without touching a file. Deterministic beats polite; a hook is only as good as its pattern, and ours matched one script.

The gate tax

Same bug, same broken start, two roads — the control experiment. Road one: just fix it. Road two: the full chain, intent to build.

direct fix full chain multiplier
Wall time 2:13 11:31 ×5.2
Cost $0.70 $3.46 ×4.9
Turns 40 169 —
Outcome suite green suite green identical

The machine bill is the small half. The chain wrote about 5,500 words of artifacts for a one-line fix — roughly 27 minutes of human reading for a diff you could scan in one. The chain converts writing time into reading time; that is the gate tax.

The playbook adds a recurring charge: continuous evals. Twenty to fifty real tasks, rerun against every config change — each case a real past task, the prompt as it was, run from the commit before the change, with checkable acceptance. Writing five cases from history took five minutes; running them right is not as easy — our first harness aimed two cases at the wrong commit, and both agents caught it instead of faking a pass. At about one minute per case, a full suite costs up to an hour of agent time per run, and every production incident is supposed to join the suite as a permanent regression eval. On a regulated team that reading is the deliverable; solo, it is overhead.

Verdict: three of six pay

Three of six gates pay for themselves:

Stage verdict evidence
Plan keep 40 s buys the questions nobody asked
Build keep plan mode + the test loop shipped 50 green tests
Deploy keep $0.80 review with real checks, 14 s deterministic block
Design skip solo bills you for policies you have not encoded
Test (continuous evals) can wait up to an hour per run, easy to aim wrong
Maintain unproven needs weeks of production telemetry

Maintain is elegant on paper — deterministic scripts watch control bands, and a breach writes a fresh intent file — but proving it needs production telemetry we do not have. Anthropic's own document, as one analyst put it, carries no measurement anywhere; these are the first numbers, with the obvious limits: one repo, one developer, one day.

The outside data says the pressure is real. Faros tracked over 10,000 developers across 1,200+ teams: high-adoption teams merge 98% more pull requests, review time rises 91%, and the average pull request more than doubles in size. The latest DORA report rhymes: throughput up with AI, stability down. Review is becoming the bottleneck, and the playbook is aimed exactly there. Community variants already cut the chain to two human decisions — one ships templates and a gate ledger, the other keeps humans at design and test only. Adopt the three gates that pay, and grow into the rest when your team does. Anthropic's own closing line is the right epitaph: the loop keeps running, human judgment stays above it.

Sources

Frequently asked questions

What is Anthropic's AI-native SDLC playbook?
A free 14-lesson Claude Academy course that structures AI development into six stages (plan, design, build, test, deploy, maintain), each ending in a committed file that the next stage reads — intent.md, spec.md, plan.md, the pull request and the incident record.
Is the AI-native SDLC playbook worth following?
Measured on a real repo, three of six gates pay for themselves: plan (29–40 s for the questions nobody asked), build (plan mode plus a TDD loop shipped a 15-file feature with 50 green tests), and deploy (an $0.80 review with real checks plus a deterministic hook block). Design, continuous evals and maintain only pay once a team encodes policies and owns production telemetry.
How much does the full artifact chain cost compared to a direct fix?
On the same bug, the direct fix took 2:13 and $0.70; the full intent → spec → plan → build chain took 11:31 and $3.46 — about five times the time and cost for an identical outcome, plus ~27 minutes of human reading.
How do Claude Code hooks work as deploy gates?
A PreToolUse hook reads every Bash command before it runs; if it matches a protected pattern (like deploy-prod), it prints a reason and exits with code 2, which blocks the tool call and hands the reason back to the model. Ours answered in 14 seconds.
What are continuous evals in the playbook?
A suite of 20–50 real past tasks rerun against every configuration change, each with checkable acceptance. Writing five cases took five minutes, but a full suite costs up to an hour of agent time per run, and every production incident is meant to join the suite as a regression eval.
Does AI coding really shift the bottleneck to review?
The field data says yes: Faros telemetry across 10,000+ developers shows high-adoption teams merge 98% more pull requests while review time rises 91% and average PR size more than doubles; the latest DORA report shows throughput up and stability down.

Related videos