TL;DR
- Playbook के छह stages में से हर एक एक committed artifact पर खत्म होता है (intent.md, spec.md, plan.md, PR, incident record)। Launch post और 14 lessons का course इसका ढांचा बताते हैं, पर कोई भी कोई माप publish नहीं करता।
- एक असली Express + Prisma repo पर शुरू से अंत तक चलाने पर, पूरी chain ने एक bug 11 min 31 s में $3.46 में ठीक किया, जबकि direct prompt ने 2 min 13 s में $0.70 में: wall time ×5.2, cost ×4.9, दोनों green।
- असली कीमत पढ़ने में है: एक लाइन के class fix के लिए 5,488 words के artifacts, 200 wpm पर करीब 27 min। Chain लिखने का समय पढ़ने के समय में बदल देती है।
- छह में से तीन stages इस context में फायदेमंद रहे: Plan (intent.md), Build (plan mode + CLAUDE.md + TDD), Deploy (REVIEW.md + एक hook)। अकेले dev के लिए Design और continuous evals फायदेमंद नहीं रहे; Maintain चलाया ही नहीं गया।
- Spec stage ने अपनी ही पूर्व शर्त बता दी: कोई org skills थे ही नहीं, इसलिए उसे brand, security या UX policy के सामने कभी जांचा नहीं गया। Playbook मानकर चलता है कि ये skills पहले से लिखे हुए हैं।
- Deterministic gate काम करता है: एक PreToolUse hook ने exit 2 के साथ deploy को 14 s में block कर दिया। Hook चलने से पहले ही model अपने judgment से एक बार मना कर चुका था।
माप क्या कहते हैं
Playbook का framing है "code अब bottleneck नहीं रहा", और वह हर stage से एक committed artifact पर खत्म होने को कहता है: intent.md से spec.md और plan.md होते हुए PR और incident record तक, Maintain में control bands के साथ s1। Course में उद्धृत करने लायक ठोस बातें हैं: eval suite के तौर पर 20 से 50 असली tasks, REVIEW.md में 5 nit की सीमा, ज़्यादा से ज़्यादा 2 से 3 parallel sessions, और यह नियम कि जो गलती दो बार हो वह CLAUDE.md में जाती है s2। सबसे तीखा निष्पक्ष विश्लेषण हर artifact के लिए यह तालिका बनाता है कि उसका draft कौन लिखता है और उसे कौन स्वीकार करता है, और दस्तावेज़ को "vendor-claim throughout" और "no measurement anywhere" कहता है s4।
Fix वाला task एक असली upstream bug था: ताज़ा clone पर npx nx test api चलाने पर शुरू से ही 1 suite fail हुई (auth.service.test.ts, "TypeError: Cannot read properties of undefined (reading 'prototype')"), 4 pass हुईं, 14 tests green, 2.2 s। Direct path 2 min 13 s, $0.70, 40 turns में पूरी तरह green tests तक पहुंचा। Chain path, यानी intent फिर spec फिर plan फिर build, भी green पर पहुंचा, 11 min 31 s, $3.46, 169 turns में। यह सिर्फ machine की तरफ से ही wall time ×5.2 और cost ×4.9 है s2।
Chain इंसान की तरफ ज़्यादा चुभती है। उसने पढ़ने के लिए 5,488 words के artifacts बनाए (intent 558 + spec 2,167 + plan 2,763), 200 wpm पर करीब 27 min, ऐसे fix के लिए जिसका direct review भार एक छोटा diff है s4। Feature task (mute authors) को पूरी chain से चलाने में 15 min 13 s, $4.11, 158 turns लगे और उसने एक Prisma Mute model migration के साथ, mute और unmute endpoints, feed filtering, 15 files में 1,422 insertions, 50 tests green, 3 नई या बढ़ाई गई test files और एक e2e spec ship किए। उसके artifacts 6,852 words के थे (intent 450 + spec 2,337 + plan 4,065), यानी करीब 34 min का पढ़ना s2।
Design stage पर संशयवादी पढ़त सही निकली। LinkedIn की आलोचना कहती है कि playbook अपनी पूर्व शर्तें छिपाता है: brand, security और UX के org skills पहले से मौजूद होने चाहिए, और किसी को पता होना चाहिए कि brainstorm कैसे चलाना है s7। Agent ने इसे बिना पूछे खुद मान लिया। spec.md की flagged concern C0 शब्दशः यह कहती है: "No org skills available. … This spec has therefore not been checked against brand, security or UX policy." 2,000 से ज़्यादा words की ऐसी spec जो codebase को दोहराती है और policy जांच नहीं सकती, अकेले काम करते समय छोड़ देने वाला stage है s7।
Infrastructure वाली आलोचना भी सही निकली। तर्क यह है कि जब tests पुराने fakes से टकराते हैं तो "the agent sees the tests pass and reports the work finished", क्योंकि artifact chain यह दर्ज करती है कि क्या तय हुआ, यह नहीं कि असल में क्या चलता है s8। इस run में loop ने सिर्फ unit tests और build जांचे; review ने खुद nx e2e को "Not run: needs a running server and a seeded DB" और prisma migrate status को "Not run: needs a DB" के रूप में सूचीबद्ध किया। Green loop ने किसी live system को छुआ ही नहीं s8।
Deploy stage सस्ती जीत थी। REVIEW.md 117 s में $0.80 में चला: nx test (5/5 suites, 50 passed), nx build (pass), plan की baseline के मुकाबले lint delta (34 बनाम 33, यह +1 plan के item A3 में साफ तौर पर अनुमत था), और एक prettier check (9 files fail, nit N1 के रूप में दर्ज)। Verdict: 0 Important, 6 nits, जिनमें 5 सूचीबद्ध और 1 सार में, क्योंकि सीमा लागू हुई। Review ने अपना ही काम approve करने से "this agent does not approve" कहकर मना कर दिया, यानी जिम्मेदारियों का अलगाव जैसा course उसे लिखता है s2। Hook gate वैसे ही चला जैसा docs बताते हैं: merge से पहले deploy करने को कहा गया तो agent ने script चलाए बिना अपने judgment से मना कर दिया, इसलिए hook कभी चला ही नहीं। Merge के बाद deploy की कोशिश को PreToolUse hook (exit 2) ने gate के संदेश के साथ 14 s में block कर दिया s19।
Evals लिखना सस्ता था और गलत करना आसान। पांच cases git history से 283 s में $1.44 में बने। दोनों runs गलत base पर चले, क्योंकि runner ने fix merge होने के बाद branch बनाई, और दोनों agents ने इसे पकड़ लिया ("the bug was already fixed here") बजाय pass का नाटक करने के। एक eval run की कीमत करीब 60 से 70 s है, इसलिए playbook के अपने 20 से 50 cases के आकार पर हर CI run में करीब 20 से 55 min agent समय लगता है s2। CLAUDE.md का setup एक committed page के लिए 63 s और $0.44 में हुआ, यानी सबसे सस्ता दांव; read only CI log triage ने सही कारण 11 s में $0.13 में पहचान लिया s2।
Community thread व्यापक telemetry लाता है: 10,000 developers में, ज़्यादा AI इस्तेमाल करने वाली teams 98% ज़्यादा PRs merge करती हैं, जबकि review का समय 91% और PR का आकार 154% बढ़ता है s6।
माप
Protocol: chain headless चली (claude -p, model claude-opus-5-5, permission acceptEdits तक सीमित और एक allowlist, --setting-sources project,local) gothinkster/node-express-realworld-example-app के scratch clone पर (Express + TypeScript + Prisma + Docker में Postgres 16, Nx workspace)। हर stage का समय नापा गया और exp/metrics.jsonl में दर्ज हुआ (17 rows)। कुल: $11.90 + hook के दोबारा run के लिए $0.14, 539 + 3 turns, करीब 41 min agent wall time।
| Stage | Wall | Turns | Cost |
|---|---|---|---|
| CLAUDE.md setup (lesson 5) | 63 s | 27 | $0.44 |
| FIX direct (बिना chain) | 133 s | 40 | $0.70 |
| FIX intent.md | 39 s | 8 | $0.22 |
| FIX spec.md | 162 s | 39 | $0.83 |
| FIX plan.md | 180 s | 46 | $1.00 |
| FIX build | 310 s | 76 | $1.40 |
| FEAT intent.md | 29 s | 6 | $0.18 |
| FEAT spec.md | 118 s | 20 | $0.62 |
| FEAT plan.md | 240 s | 41 | $1.17 |
| FEAT build (TDD) | 526 s | 91 | $2.14 |
| Review (REVIEW.md) | 117 s | 19 | $0.80 |
| Hook demo (मना किया) | 20 s | 5 | $0.14 |
| Hook demo (block हुआ) | 14 s | 3 | $0.14 |
| CI triage (read only) | 11 s | 3 | $0.13 |
| Evals: 5 cases लिखना | 283 s | 76 | $1.44 |
| Eval run 1 / run 2 | 72 s / 59 s | 24 / 18 | $0.38 / $0.29 |
| Playbook stage | Verdict | कारण |
|---|---|---|
| Plan (intent.md) | रखें | 29 से 39 s, असली खुले सवाल सामने लाता है, चुपचाप होने वाले architecture चुनाव खत्म करता है |
| Design (spec.md) | अकेले हों तो छोड़ें | 2,000 से ज़्यादा words जो codebase को दोहराते हैं; इसका मूल्य उन org skills पर टिका है जो मौजूद नहीं हैं (इसका अपना C0 flag) |
| Build (plan mode + CLAUDE.md + TDD loop) | रखें | 50 tests green, deviations दर्ज, review plan पर टिका रहा |
| Test (continuous evals) | फिलहाल छोड़ें | playbook के अपने आकार पर हर CI run में 20 से 55 min; base commit का अनुशासन सबसे पहले टूटा |
| Deploy (REVIEW.md + hooks) | रखें | असली checks वाला $0.80 का review और 14 s का deterministic block |
| Maintain (control bands) | सिद्ध नहीं | हफ्तों की production telemetry चाहिए; अनुमानित, चलाया नहीं गया |
सीमाएं: एक repo, एक developer, एक दिन। Team स्तर के plays नहीं चलाए गए, headless mode interview के steps को एक ही prompt में समेट देता है, और eval runs की गिनती सिर्फ प्रति run cost के लिए है।
सोमवार को यह करें
- अपनी मुख्य repo के लिए CLAUDE.md का एक page लिखें: build, test और lint commands, और पिछले हफ्ते agent की की हुई दो गलतियां। उसे commit करें। Agent समय का बजट 63 s।
- अपने अगले non trivial task से पहले सबसे पहले intent.md मांगें: goal, non goals, खुले फैसले। खुले सवालों के जवाब दें, फिर agent को plan बनाने दें। जब तक आपके पास spec जांचने के लिए org policy skills न हों, spec.md छोड़ दें।
- Build stage को plan mode में TDD loop के साथ चलाएं और plan से deviations (D1, D2, ...) दर्ज करवाएं ताकि review के पास टिकने के लिए कुछ हो।
- एक नए session से REVIEW.md pass चलवाएं, जिसमें nit की सीमा हो और साफ "this agent does not approve" पंक्ति हो। उससे tests, build, lint delta और formatter check चलवाएं।
- एक deterministic gate लगाएं: ऐसा PreToolUse hook जो branch main न होने पर
deployपर exit 2 करे। - Green loop पर भरोसा करने से पहले, review के अंत में लिखें कि उसने क्या नहीं चलाया (e2e, migrations, जिसे भी live DB चाहिए)।
- अपना gate tax खुद नापें: एक ही छोटे bug पर direct path और chain path का समय लें, फिर गिनें कि आपको कितने words पढ़ने पड़े।
आगे बढ़ें
- दो gate वाला variant: एक adversarial review gate (sdlc-gate) और हर stage पर एक के बजाय सिर्फ दो human decision points, छोटी team के लिए व्यावहारिक रूप s12।
- पूरी installable chain: intent, spec, plan और REVIEW के templates, एक gate validator, एक eval runner और control band detection, अगर आप scaffolding हाथ से नहीं बनाना चाहते s5।
- Interview से शुरू होने वाली planning: एक बार में एक सवाल, सवालों को batch करने से बेहतर है, और "AI agents don't ask clarifying questions. They assume." समय के आंकड़ों के बिना एक setup write up s11।
- एक तय pipeline को लोग क्यों घुमाकर किनारे कर देते हैं: "a docs fix and a payments migration shouldn't travel the same path", और असली process अदृश्य हो जाता है। Playbook को Kiro और GitHub Spec Kit के साथ रखता है s9।
- वे खाली जगहें जो कोई platform vendor आपको बेचेगा: signal से intent तक का intake, blast radius routing, एक metrics dashboard s13।
- एक consultancy जो जनवरी से client teams पर वही ढांचा (CRAFT) चला रही है और मानती है कि "we don't yet have a formal answer for what a control band looks like" s10।
- एक काम करता हुआ intent.md उदाहरण (Select All checkbox) जो फाइल का काम दिखाता है: agent को चुपचाप चुनने देने के बजाय खुले फैसले सामने लाना s14।
स्रोत
- The AI-Native SDLC Playbook (launch post), claude.com. यह क्यों पढ़ें: छह stages का ढांचा और committed artifact का नियम पांच मिनट में।
- The AI-native SDLC playbook (course, 14 lessons), Claude Academy. यह क्यों पढ़ें: एकमात्र जगह जहां संख्याएं मिलती हैं (20 से 50 eval tasks, 5 nit की सीमा, 2 से 3 sessions), मुफ्त और बिना login।
- The Committed-Artifact Chain, howardism.dev. यह क्यों पढ़ें: हर artifact का draft कौन लिखता है और कौन स्वीकार करता है, और यह सीधा निष्कर्ष कि playbook में कुछ भी नापा नहीं गया।
- bashebr/ai-native-sdlc, GitHub. यह क्यों पढ़ें: templates, gate validator और eval runner जिन्हें आप खुद लिखने के बजाय install कर सकते हैं।
- Anthropic published an AI-native SDLC playbook, r/ClaudeAI. यह क्यों पढ़ें: वह thread जो Faros की telemetry (98% ज़्यादा PRs, +91% review समय) को बहस में लाता है।
- The AI-native SDLC Playbook is basically "do everything you did before, but inside Claude", LinkedIn. यह क्यों पढ़ें: छिपी पूर्व शर्तों का तर्क, जिसे spec stage ने खुद सही ठहराया।
- The AI-Native SDLC Starts With Your Infrastructure, MetalBear blog. यह क्यों पढ़ें: पुराने fakes की समस्या; vendor की आवाज़ है, पर तर्क अपने दम पर खड़ा है।
- The AI-native SDLC won't be one process, worldprogramming.org. यह क्यों पढ़ें: हर बदलाव के लिए एक ही रास्ते के खिलाफ औपचारिकता का तर्क।
- Anthropic Wrote the AI-Native SDLC Playbook in August. We Wrote Ours in January., Substack. यह क्यों पढ़ें: एक स्वतंत्र team जो उसी ढांचे पर पहुंची और Maintain की कमी मानती है।
- AI-Native SDLC: First Try, kyle.pericak.com. यह क्यों पढ़ें: interview से शुरू होने वाला एकमात्र hands on पहला run, playbook के आने से पहले लिखा गया।
- TsCarpe/claude-sdlc-skills, GitHub. यह क्यों पढ़ें: adversarial review step वाला दो gate का variant।
- Implementing the Anthropic AI-Native SDLC Playbook, Port blog. यह क्यों पढ़ें: उन चीज़ों की सूची जो playbook छोड़ देता है, खाली जगहों के नक्शे की तरह पढ़ें।
- What Is intent.md in Claude Code?, dev.to. यह क्यों पढ़ें: एक ठोस intent.md जिसकी संरचना आप नकल कर सकते हैं।
- Hooks guide, Claude Code docs. यह क्यों पढ़ें: exit 2 वाला PreToolUse hook कैसे वह deterministic gate बनता है जिस पर playbook निर्भर है।
FAQ
क्या एक लाइन के fix के लिए पूरी chain कभी फायदेमंद है?
इस run में नहीं: उसी green नतीजे के लिए wall time ×5.2 और cost ×4.9, ऊपर से पढ़ने के लिए 5,488 words। छोटे tasks के लिए सिर्फ intent.md इस्तेमाल करें।
अकेले काम करते समय spec.md क्यों छोड़ें?
Spec ने खुद बता दिया: brand, security या UX के org skills न होने से वह policy जांच नहीं सकी, और उसने 2,000 से ज़्यादा words codebase को दोहराने में खर्च किए।
क्या hook model के judgment की जगह लेता है?
नहीं, वह उसका साथ देता है। Agent ने merge से पहले का deploy खुद मना कर दिया; hook ने merge के बाद की कोशिश को exit 2 के साथ 14 s में block किया।
AIDive