TL;DR
- On a real repo with a 177-line CLAUDE.md, 3 skills and 1 hook, deleting everything broke exactly one rule, in one situation: the i18n rule on a brand-new file. Every other convention held because the surrounding code already taught it.
- The file cost 4 to 14% of the tokens read on tasks where the output was identical, about a dollar in ten. The headline 32% saving is dominated by the one task where the file made the model do more work.
- Skills and the hook cost nothing measurable: skills load on invoke and none was invoked, the hook injects one sentence.
- The 82-line architecture overview bought nothing on the architecture question: six answers, all correct, with or without the file.
- "Delete" in Anthropic's own words means ablate and move to progressive disclosure, not erase. Keep the rules the code cannot teach, move the rest to rules files and skills, and turn the non-negotiables into hooks.
- Two reps per config is a floor, not a verdict: per-task deltas under 15% are within run-to-run noise.
What the measurements say
The talk everyone reacts to says two things, and the second gets dropped. Boris Cherny confirms on stage that Claude Code deleted 80% of its system prompt, and he describes the method: delete the entire system prompt, then bring it back line by line to measure the impact of each line s2. The "every 6 months" passage sits at 00:06:58 to 00:07:05, and the wording is "really do recommend", not "strongly recommend" s1. Two lines widely attributed to him are not in the talk: "context, goals and a definition of done" (the closest real phrase at 15:22 is "describe the task, the guardrails, the exit criteria") and "64 agents". He says "eleven days" and, asked for an agent count, "I'm not sure" s2. The 64 figure belongs to the Bun rewrite post: "64 Claudes running for 11 days", around $165,000 at API pricing, 9 billion uncached input tokens, 690 million output tokens and 72 billion cached input token reads s14.
The mechanism that makes any of this measurable is loading. CLAUDE.md and rules files are injected every turn; skills load only when invoked. The memory docs also carry the size guidance everyone paraphrases: "target under 200 lines per CLAUDE.md file. Longer files consume more context and reduce adherence." s4. Anthropic's written version of the advice frames the central-repository CLAUDE.md as a myth and points to progressive disclosure and auto-memory instead s3.
The only controlled study before ours is the ETH paper on AGENTS.md: 138 issues across 12 repos, and "providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average". Developer-committed files outperform LLM-generated ones by 7% on average, and instructions that name specific tools do help s5.
Two data points on what the files contain and how they are read. Across 28,721 repositories and 165,063 files, the median instruction file holds 50 content items, of which 12 are actual directives; the author's framing is that only 27% of the file is doing what you think it does s8. In a controlled experiment with two instructions in genuine conflict, moving a single rule from the top of the file to the bottom swings how often the model obeys it by about 90 points, from almost never to almost always s9. The same publisher draws the line our experiment follows: ablation is not deletion, and hooks are enforcement, they do not expire with a model upgrade s7.
Two informal priors matched our result. A 1,000-line CLAUDE.md compared against a stripped ~20-line version on the same GitHub issues with the same model: architecture held, house rules broke s6. A commenter who moved everything to progressive disclosure reports auto-loaded context dropping from ~35k tokens to ~4.5k per session, with zero knowledge deleted s11. For scoring skill ablations there is a tool: caliper's --ablate covers skills and MCP servers and its compare reports success rate, tokens and wall time; the CLAUDE.md swap stays manual s16.
Measurements
Protocol: one private Expo / React Native repo (1,021 tracked files), CLAUDE.md of 177 lines, 7,243 characters, about 1,800 tokens, 3 skills, 4 slash commands, 1 PreToolUse hook. Model pinned to claude-opus-5 for every run, Claude Code 2.1.278, headless claude -p, --max-turns 40, user config directory empty, no MCP, fresh clone reset before each run. Five everyday tasks (new component, edit a component, refactor shared helpers, persist a store field, explain a data path), ablation one piece at a time. 44 runs, $38.97 at API pricing, 92 minutes of agent wall time. Scoring: the repo's own house rules on added lines, tsc --noEmit, eslint, answers graded by hand.
Quality, what broke:
| config | edit runs | house-rule violations | tsc new errors | eslint errors |
|---|---|---|---|---|
| A full | 8 | 0 | 0 | 0 |
| B no CLAUDE.md | 8 | 1 (new heading hardcoded, no .content.ts) |
0 | 0 |
| C no skills | 4 | 0 | 0 | 0 |
| D no hook | 4 | 0 | 0 | 0 |
| E nothing | 8 | 2 (heading hardcoded twice, no .content.ts) |
0 | 0 |
Cost per run, mean over the config's runs (tokens = cache reads, the context re-read on every turn):
| config | runs | $ / run | turns / run | tokens read / run |
|---|---|---|---|---|
| A full | 10 | 0.95 | 25.6 | 829k |
| B no CLAUDE.md | 10 | 0.74 | 21.9 | 568k |
| C no skills | 5 | 0.96 | 28.2 | 819k |
| D no hook | 5 | 0.96 | 27.8 | 851k |
| E nothing | 10 | 0.85 | 25.7 | 655k |
Per task, A to B (mean of 2 runs each):
| task | A $ | B $ | cost | tokens read |
|---|---|---|---|---|
| T1 new component | 1.72 | 0.96 | -44% | -61% |
| T2 edit component | 1.49 | 1.26 | -15% | -15% |
| T3 refactor | 0.64 | 0.63 | -1% | -5% |
| T4 store | 0.57 | 0.53 | -6% | -4% |
| T5 question | 0.34 | 0.31 | -9% | -14% |
| all 10 runs | 9.51 | 7.40 | -22% | -32% |
Reading the tables. Without the CLAUDE.md, 3 of 4 new-component runs wrote the heading as a plain string; with it, 4 of 4 created the .content.ts with EN and FR. On the edit task every config, even E, put the new string in a .content.ts: the neighbouring files carried the convention. Everything else held in all 44 runs: 0 relative imports, type not interface in six edits of the types file, design tokens in every diff, 0 hex colours, no barrel file. One instruction nobody could follow: the file says to import theme from @design-tokens, an alias that does not exist in tsconfig.json; no run in any config used it, 1,800 tokens read for nothing every session. The T1 saving is the cheaper run doing less work. Run-to-run variance is larger than most config effects: T1 with the full config cost $2.11 then $1.33. A control with the 333 MB code graph present landed on the full config's numbers ($1.59 vs $1.72 on T1, $0.38 vs $0.34 on T5): the graph block and the hook changed nothing measurable.
Do this Monday
- Count your CLAUDE.md: lines, characters, and the number of lines that are actual directives (Always, Never, Use, Prefer). The rest is context the model re-reads every turn.
- Pick one convention per section and check whether a neighbouring file already shows it. If three files in the folder follow the rule, the line is a candidate for deletion.
- Find the one rule the code cannot teach on a brand-new file (i18n, telemetry, licence headers, a required registration step). Keep it, state it in one line, and put it near the top.
- Try every import path and command your file names. A dead alias or a renamed script is an instruction nobody can follow.
- Move long architecture overviews and setup walkthroughs to a rules file or a skill that loads on demand, then compare the auto-loaded context before and after.
- Turn your two or three non-negotiables into a hook or a lint rule. Enforcement does not depend on the model reading a paragraph.
- Run your two most common tasks twice with the file and twice without, on a pinned model, and read tokens and the diff. Two reps tell you where to look, not what to conclude.
Go further
- The talk itself, for the method rather than the soundbite: delete, then bring back line by line s2.
- The paper's full result set, including the finding that developer-committed files beat generated ones by 7% and that naming tools helps s5.
- Rule position as a variable: the ~90-point swing and how the model resolves conflicting instructions silently s9.
- The 30k-repo anatomy of an instruction file: 50 items, 12 directives, and what the other 38 are s8.
- The HumanLayer guide to writing a good CLAUDE.md and the thread that argues with it s10.
- The "MUST use agent, ignored 80% of the time" thread: a case for hooks where prose fails s12.
- The "Claude Code now ignores everything" thread, useful for separating model drift from instruction conflicts s13.
- caliper, to score skill and MCP ablations with success rate, tokens and wall time s16.
Sources
- Boris Cherny: We Cut 80% of Claude Code's Prompt, Y Combinator. Why read it: the primary quote, with the caveat the clips cut.
- Transcript of the talk, Y Combinator. Why read it: searchable text to check what was and was not said.
- The new rules of context engineering for Claude 5 generation models, Anthropic. Why read it: the written version of the advice, progressive disclosure over a central file.
- Memory docs: CLAUDE.md and rules files, Claude Code docs. Why read it: what loads every turn, what loads on demand, and the 200-line guidance.
- Evaluating AGENTS.md, arXiv. Why read it: the only controlled measurement of context files before this one.
- Should you delete your CLAUDE.md every 6 months?, Modern Creator. Why read it: an informal 1,000-line vs 20-line comparison whose breakage pattern matches ours.
- Opus 5: delete your CLAUDE.md?, reporails. Why read it: ablation versus deletion, and why hooks survive a model upgrade.
- The State of AI Instruction Quality: 30k-repo analysis, reporails. Why read it: what a median instruction file actually contains.
- Opus 5: the cost of instruction conflicts, reporails. Why read it: a controlled experiment on rule position.
- Writing a good CLAUDE.md, HN thread, Hacker News. Why read it: practitioners arguing over what belongs in the file.
- Anthropic says keep CLAUDE.md under 200 lines, r/ClaudeCode. Why read it: the ~35k to ~4.5k before/after comment.
- CLAUDE.md says MUST use agent, Claude ignores it, r/ClaudeCode. Why read it: a concrete case where prose instructions fail.
- Claude Code now ignores everything, r/ClaudeCode. Why read it: the symptom reports behind the "something changed" feeling.
- Rewriting Bun in Rust, Simon Willison. Why read it: where the 64 agents and $165,000 figures come from.
- caliper, GitHub. Why read it: a scoring harness for skill and MCP ablations.
FAQ
Should I delete my CLAUDE.md?
Not blindly. On our repo the only loss was one rule on new files; everything the code already demonstrated held without the file. Ablate one section at a time and keep what changes the output.
Why did the file save 32% of tokens if it only cost 4 to 14%?
Because the aggregate is dominated by the new-component task, where the file made the model write a second file in two languages. The cheaper run did less. On tasks with identical output, the saving was 4 to 14%.
Do skills and hooks cost tokens every turn?
Skills load only when invoked, so an uninvoked skill costs nothing; the hook injects one sentence. Deleting both changed no number in our runs. Rules files and CLAUDE.md are the ones re-read every turn.
Is two runs per configuration enough?
No. Two reps locate the effect; they do not size it. Our full config cost $2.11 then $1.33 on the same task, so per-task deltas under 15% are within noise.
AIDive