TL;DR
- The playbook's six stages each end in a committed artifact (intent.md, spec.md, plan.md, PR, incident record). The launch post and the 14 lesson course describe the shape; neither publishes a measurement.
- Run end to end on a real Express + Prisma repo, the full chain fixed one bug in 11 min 31 s for $3.46 where a direct prompt did it in 2 min 13 s for $0.70: ×5.2 wall, ×4.9 cost, both green.
- The real tax is reading: 5,488 words of artifacts for a one line class fix, about 27 min at 200 wpm. The chain converts writing time into reading time.
- Three of six stages paid in this context: Plan (intent.md), Build (plan mode + CLAUDE.md + TDD), Deploy (REVIEW.md + a hook). Design and continuous evals did not for a solo dev; Maintain was not run.
- The spec stage flagged its own prerequisite: no org skills existed, so it was never checked against brand, security or UX policy. The playbook assumes those skills are already written.
- The deterministic gate works: a PreToolUse hook blocked a deploy in 14 s with exit 2. The model had already refused once on its own judgment before the hook ever fired.
What the measurements say
The playbook frames the shift as "code is no longer the bottleneck" and asks every stage to end in a committed artifact, from intent.md through spec.md and plan.md to the PR and the incident record, with control bands in Maintain s1. The course carries the quotable specifics: 20 to 50 real tasks as the eval suite, a 5 nit cap in REVIEW.md, 2 to 3 parallel sessions at most, and the rule that a mistake made twice goes into CLAUDE.md s2. The sharpest neutral read tabulates who drafts and who accepts each artifact and calls the document "vendor-claim throughout" with "no measurement anywhere" s4.
The fix task was a real upstream bug: a fresh clone ran npx nx test api and 1 suite failed out of the box (auth.service.test.ts, "TypeError: Cannot read properties of undefined (reading 'prototype')"), 4 passed, 14 tests green, 2.2 s.The direct path reached fully green tests in 2 min 13 s, $0.70, 40 turns. The chain path, intent then spec then plan then build, also reached green in 11 min 31 s, $3.46, 169 turns. That is ×5.2 wall and ×4.9 cost on the machine side alone s2.
The human side is where the chain hurts. It produced 5,488 words of artifacts to read (intent 558 + spec 2,167 + plan 2,763), about 27 min at 200 wpm, for a fix whose direct review load is one small diff s4. The feature task (mute authors) through the full chain took 15 min 13 s, $4.11, 158 turns and shipped a Prisma Mute model with migration, mute and unmute endpoints, feed filtering, 1,422 insertions over 15 files, 50 tests green with 3 new or extended test files and an e2e spec. Its artifacts weighed 6,852 words (intent 450 + spec 2,337 + plan 4,065), roughly 34 min of reading s2.
The skeptical reading of the Design stage held up. The LinkedIn critique says the playbook hides its prerequisites: org skills for brand, security and UX must already exist, and someone must know how to run the brainstorm s7. The agent confirmed it unprompted. spec.md's flagged concern C0 reads, verbatim: "No org skills available. … This spec has therefore not been checked against brand, security or UX policy." A 2,000 plus word spec that restates the codebase and cannot check policy is the stage to skip when working alone s7.
The infrastructure critique held too. The argument is that when tests hit stale fakes "the agent sees the tests pass and reports the work finished", because the artifact chain records what was decided, not what actually runs s8. In the run, the loop verified unit tests and the build only; the review itself listed nx e2e as "Not run: needs a running server and a seeded DB" and prisma migrate status as "Not run: needs a DB". The green loop never touched a live system s8.
The Deploy stage was the cheap win. REVIEW.md ran in 117 s for $0.80: nx test (5/5 suites, 50 passed), nx build (pass), a lint delta against the plan baseline (34 vs 33, the +1 explicitly allowed by plan item A3), and a prettier check (9 files failing, logged as nit N1). Verdict: 0 Important, 6 nits, 5 listed and 1 summarized because the cap applied. The review refused to approve its own work with "this agent does not approve", the separation of duties as the course writes it s2. The hook gate behaved as the docs describe: asked to deploy before the merge, the agent refused on its own judgment without running the script, so the hook never fired. After the merge the deploy attempt was blocked by the PreToolUse hook (exit 2) in 14 s with the gate's message s19.
Evals were cheap to write and easy to get wrong. Five cases came out of git history in 283 s for $1.44. Both runs executed against the wrong base, because the runner branched after the fix was merged, and both agents detected it ("the bug was already fixed here") instead of faking a pass. One eval run costs about 60 to 70 s, so the playbook's own sizing of 20 to 50 cases means roughly 20 to 55 min of agent time per CI run s2. CLAUDE.md setup took 63 s and $0.44 for one committed page, the cheapest play of all; a read only CI log triage named the correct cause in 11 s for $0.13 s2.
The community thread imports the wider telemetry: across 10,000 developers, high AI teams merge 98% more PRs while review time rises 91% and PR size 154% s6.
Measurements
Protocol: the chain ran headless (claude -p, model claude-opus-5-5, permission scoped acceptEdits plus an allowlist, --setting-sources project,local) on a scratch clone of gothinkster/node-express-realworld-example-app (Express + TypeScript + Prisma + Postgres 16 in Docker, Nx workspace). Every stage was timed and logged to exp/metrics.jsonl (17 rows). Total: $11.90 + $0.14 for a hook rerun, 539 + 3 turns, about 41 min of agent wall time.
| Stage | Wall | Turns | Cost |
|---|---|---|---|
| CLAUDE.md setup (lesson 5) | 63 s | 27 | $0.44 |
| FIX direct (no chain) | 133 s | 40 | $0.70 |
| FIX intent.md | 39 s | 8 | $0.22 |
| FIX spec.md | 162 s | 39 | $0.83 |
| FIX plan.md | 180 s | 46 | $1.00 |
| FIX build | 310 s | 76 | $1.40 |
| FEAT intent.md | 29 s | 6 | $0.18 |
| FEAT spec.md | 118 s | 20 | $0.62 |
| FEAT plan.md | 240 s | 41 | $1.17 |
| FEAT build (TDD) | 526 s | 91 | $2.14 |
| Review (REVIEW.md) | 117 s | 19 | $0.80 |
| Hook demo (refused) | 20 s | 5 | $0.14 |
| Hook demo (blocked) | 14 s | 3 | $0.14 |
| CI triage (read only) | 11 s | 3 | $0.13 |
| Evals: write 5 cases | 283 s | 76 | $1.44 |
| Eval run 1 / run 2 | 72 s / 59 s | 24 / 18 | $0.38 / $0.29 |
| Playbook stage | Verdict | Why |
|---|---|---|
| Plan (intent.md) | Keep | 29 to 39 s, surfaces real open questions, kills silent architecture picks |
| Design (spec.md) | Skip solo | 2,000 plus words restating the codebase; its value assumes org skills that do not exist (its own C0 flag) |
| Build (plan mode + CLAUDE.md + TDD loop) | Keep | 50 tests green, deviations recorded, the review leaned on the plan |
| Test (continuous evals) | Skip for now | 20 to 55 min per CI run at the playbook's own sizing; base commit discipline failed first |
| Deploy (REVIEW.md + hooks) | Keep | $0.80 review with real checks plus a deterministic 14 s block |
| Maintain (control bands) | Not proven | needs weeks of production telemetry; projected, not run |
Caveats: one repo, one developer, one day. Team scale plays were not run, headless mode compresses the interview steps into single prompts, and eval runs count for per run cost only.
Do this Monday
- Write one page of CLAUDE.md for your main repo: build, test and lint commands, the two mistakes the agent made last week. Commit it. Budget 63 s of agent time.
- Before your next non trivial task, ask for an intent.md first: goal, non goals, open decisions. Answer the open questions, then let the agent plan. Skip spec.md unless you have org policy skills to check it against.
- Run the build stage in plan mode with a TDD loop and require the plan to record deviations (D1, D2, ...) so the review has something to lean on.
- Add a REVIEW.md pass run by a fresh session with a nit cap and an explicit "this agent does not approve" line. Have it run tests, build, a lint delta and a formatter check.
- Put one deterministic gate in place: a PreToolUse hook that exits 2 on
deploywhen the branch is not main. - Before trusting a green loop, list what it did not run (e2e, migrations, anything needing a live DB) at the bottom of the review.
- Measure your own gate tax: time the direct path and the chain path on the same small bug, then count the words you had to read.
Go further
- The two gate variant: an adversarial review gate (sdlc-gate) and only two human decision points instead of one per stage, the pragmatic shape for a small team s12.
- The complete installable chain: intent, spec, plan and REVIEW templates, a gate validator, an eval runner and control band detection, if you would rather not hand build the scaffolding s5.
- Interview first planning: one question at a time beats batching, and "AI agents don't ask clarifying questions. They assume." A setup write up with no timings s11.
- Why one fixed pipeline gets routed around: "a docs fix and a payments migration shouldn't travel the same path", and the real process becomes invisible. Groups the playbook with Kiro and GitHub Spec Kit s9.
- The gaps a platform vendor will sell you: signal to intent intake, blast radius routing, a metrics dashboard. s13.
- A consultancy that ran the same shape (CRAFT) at client teams since January and admits "we don't yet have a formal answer for what a control band looks like" s10.
- One worked intent.md example (a Select All checkbox) that shows the file's job: surface open decisions instead of letting the agent pick silently s14.
Sources
- The AI-Native SDLC Playbook (launch post), claude.com. Why read it: the six stage shape and the committed artifact rule in five minutes.
- The AI-native SDLC playbook (course, 14 lessons), Claude Academy. Why read it: the only place the numbers live (20 to 50 eval tasks, 5 nit cap, 2 to 3 sessions), free and no login.
- The Committed-Artifact Chain, howardism.dev. Why read it: who drafts and who accepts each artifact, and the blunt call that nothing in the playbook is measured.
- bashebr/ai-native-sdlc, GitHub. Why read it: templates, gate validator and eval runner you can install instead of writing.
- Anthropic published an AI-native SDLC playbook, r/ClaudeAI. Why read it: the thread that brings the Faros telemetry (98% more PRs, +91% review time) into the debate.
- The AI-native SDLC Playbook is basically "do everything you did before, but inside Claude", LinkedIn. Why read it: the hidden prerequisites argument that the spec stage confirmed on its own.
- The AI-Native SDLC Starts With Your Infrastructure, MetalBear blog. Why read it: the stale fakes problem, vendor voiced but the argument stands on its own.
- The AI-native SDLC won't be one process, worldprogramming.org. Why read it: the ceremony argument against one path for every change.
- Anthropic Wrote the AI-Native SDLC Playbook in August. We Wrote Ours in January., Substack. Why read it: an independent team converging on the same shape and admitting the Maintain hole.
- AI-Native SDLC: First Try, kyle.pericak.com. Why read it: the only hands on interview first run, written before the playbook existed.
- TsCarpe/claude-sdlc-skills, GitHub. Why read it: the two gate variant with an adversarial review step.
- Implementing the Anthropic AI-Native SDLC Playbook, Port blog. Why read it: the list of what the playbook leaves out, read as a map of gaps.
- What Is intent.md in Claude Code?, dev.to. Why read it: one concrete intent.md you can copy the structure of.
- Hooks guide, Claude Code docs. Why read it: how a PreToolUse hook with exit 2 becomes the deterministic gate the playbook relies on.
FAQ
Is the full chain ever worth it for a one line fix?
Not in this run: ×5.2 wall and ×4.9 cost for the same green result, plus 5,488 words to read. Use intent.md alone for small tasks.
Why skip spec.md when working solo?
The spec flagged itself: with no org skills for brand, security or UX it could not check policy, and it spent 2,000 plus words restating the codebase.
Does the hook replace the model's judgment?
No, it backs it. The agent refused the pre merge deploy on its own; the hook blocked the post merge attempt in 14 s with exit 2.
AIDive