TL;DR
- Spotify's "90%" is the mean of bulk-read scenarios measured in estimated input tokens on a Java monorepo. The post gives no dollar figure and no quality score.
- Rebuilt in plain Claude Code (one PreToolUse hook, two cheap subagents, a three-line routing rule) and measured on Fastify over four scenarios and 16 runs, the pattern cut the main model's context by 59.6% and total cost by 33.1%.
- The deny hooks fired zero times in the measured runs. The savings came from the routing rule in CLAUDE.md; the hooks are the safety net for the day the model ignores it.
- Delegation was slower every time, +65.3% wall time on average. On the small test-writing scenario it cost 2.6% more.
- Two traps: hooks also fire inside subagents, so exempt your workers, and
sed -nrange reads pass straight through a hook that only watchescat,headandtail. - The Haiku reader's summary carried factual errors in two of eight delegated runs. Keep the main model's verification turn.
What the measurements say
Spotify's plugin, Shunt, routes bulk work away from the main model through two "modes": a bulk-reader and a code-writer, both running Gemini 2.5 Flash in the examples, with the model field accepting any model configured in the Portal instance s1. The routing has three layers. A check-file-size hook fires on every Read and blocks files over a configurable line threshold, default 350, pointing the model to the bulk-reader skill instead; a check-bash-read hook catches cat, head, tail, less and more on large files while piped commands pass through s1. The hook sources and the two skills are in the public repo s2, with the size check readable on its own s3. The modes themselves live in Portal, Spotify's internal platform, which is why the plugin cannot be run outside the company as shipped s4.
The benchmark claim is thin. Spotify tested four scenarios on a Java monorepo, "measuring tokens Claude would consume reading files directly vs. consuming the bulk-reader's summary", and reports mean bulk-read savings of around 90% s1. The post itself says the code-write scenario is harder to measure in tokens, that worker summaries do not include reliable line numbers so editing cannot be delegated, that the worker missed a subtle thread-safety bug the main model caught, and that each delegation adds 10 to 30 seconds with Portal capping a single invocation at 30 seconds s1. The Hacker News thread raised the same questions about what the 90% measures s7.
The rebuild replaces Portal modes with two Claude Code subagents whose definition files pin the model: an Explore reader on Haiku and a code-writer on Sonnet s6. The deny is a PreToolUse hook returning the deny decision in the current hook JSON format s5. The repo under test was fastify/fastify at commit ac28821d, 294 .js/.ts files, 78270 lines, 63 files over 350 lines. Model ids as reported by the session JSON: main conversation claude-opus-5[1m], reader claude-haiku-4-5-20251001, writer claude-sonnet-5. Each scenario ran twice per config, 16 measured runs, single-turn claude -p sessions with project-only settings so both sides had an identical system prompt s5.
Where the pattern won: S2, a three-file call-graph question over lib/route.js (691 lines), lib/reply.js (1090) and lib/request.js (398), went from a mean main context of 357165.5 tokens to 73440.0 (-79.4%) and from 0.5810500000000001 USD to 0.21823605000000001 USD (-62.4%). Where it did not: S4, writing a test for a 45-line source against a 19-line reference, cost 0.29465575 USD without delegation and 0.3022213 USD with it (+2.6%), because Sonnet is a second full context (13004 to 18729 cache-read tokens) and the main model still re-read the generated file and ran the test s6.
Three findings matter more than the percentages. First, the hooks fired zero times across the 16 measured runs: with the routing rule in CLAUDE.md, the main model ran wc -l and delegated on its own. The only observed deny came from a verification run without CLAUDE.md, where the model was refused Read, then refused cat -n, and answered from grep -n alone without ever calling the Agent tool s5. Second, hooks run inside subagents: in two discarded runs the Haiku reader was itself denied by the size check and fell back to chunked offset/limit reads. The fix is a case "$agent_type" in Explore|code-writer) exit 0 escape at the top of the hook, with the field name confirmed from the logged stdin s5. Third, in the baseline config the main model never used the Read tool at all. It read every file through Bash (cat -n, sed -n '1,200p', sed -n '200,560p'), so a hook that only watches Read catches nothing, and a Bash hook that only matches unpiped cat, head and tail still lets sed -n ranges through s3.
Quality was checked against the source with grep. The baseline produced wrong line numbers in one S2 run (it dumped files with sed -n without line numbers and counted by hand). The delegated config produced three factual errors in the other S2 run and two in an S3 run, all traced to the Haiku summary being taken at face value: wrong callers for buildRequest/buildReply, a non-exported constant listed as an export, a covered iterator marked uncovered. Where the main model spent output tokens re-verifying with grep (S3, 4534 to 4738 output tokens), the answers held s6. On the test-writing scenario every generated file passed: 12/12, 6/6, 7/7 and 10/10 tests. One Reddit report shows the related failure mode of a main model spawning workers on the wrong model when nothing pins it s8, which is what the require-model hook and the model: field in the agent files guard against.
Measurements
Means of the 2 runs per cell. "Main context" is input + cache_creation + cache_read tokens billed to the main model over the session, the figure comparable to Spotify's "tokens in the main context". A = plain Claude Code, B = hook + subagents + CLAUDE.md rule.
| scenario | main context A | main context B | change | main output A | main output B | change | total cost A | total cost B | change | duration A s | duration B s | change |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| S1 | 88693.0 | 51551.5 | -41.9% | 1424.5 | 1060.5 | -25.6% | 0.13910675 | 0.08653685 | -37.8% | 21.817500000000003 | 43.799499999999995 | +100.8% |
| S2 | 357165.5 | 73440.0 | -79.4% | 3835.0 | 2555.5 | -33.4% | 0.5810500000000001 | 0.21823605000000001 | -62.4% | 51.637 | 129.036 | +149.9% |
| S3 | 303807.5 | 114135.5 | -62.4% | 6192.0 | 4636.0 | -25.1% | 0.451037 | 0.3738534 | -17.1% | 93.321 | 124.64099999999999 | +33.6% |
| S4 | 143431.5 | 121818.0 | -15.1% | 5275.5 | 3340.0 | -36.7% | 0.29465575 | 0.3022213 | +2.6% | 65.7125 | 86.857 | +32.2% |
| all 4 | 223274.375 | 90236.25 | -59.6% | 4181.75 | 2898.0 | -30.7% | 0.366462375 | 0.2452119 | -33.1% | 58.122 | 96.08337499999999 | +65.3% |
Protocol: two byte-identical shallow clones of fastify/fastify at ac28821d; repo-shunt adds .claude/ (settings, two agent files, three hooks) and a CLAUDE.md routing rule, nothing else. Every session: claude -p "<prompt>" --output-format json --setting-sources project --strict-mcp-config with an empty MCP config, no --model, 600 s timeout. Four prompts, identical on both sides: S1 exports of lib/reply.js, S2 call graph across three lib files, S3 methods of lib/hooks.js versus test/hooks.test.js coverage, S4 write test/head-route.test.js following test/noop-set.test.js. Figures are read from the session JSON modelUsage and total_cost_usd, unrounded. Answers were checked against the source with grep; generated tests were run with node --test.
Do this Monday
- Run
wc -lover your repo and count files over 350 lines. If the count is near zero, stop here: the threshold exists because delegation costs more than it saves on small files. - Add a three-line routing rule to your CLAUDE.md: files over the threshold go to a reader subagent, pattern-following code goes to a writer subagent, debugging and architecture stay with the main model. In the measurements this rule did all the work.
- Create
.claude/agents/Explore.mdwithmodel: haikuand.claude/agents/code-writer.mdwithmodel: sonnetin the frontmatter, so the worker model is pinned in the file and not left to the orchestrator. - Write the PreToolUse hook on Read as a safety net, returning the deny decision in the current hook JSON format, and make its first lines exit 0 when
agent_typeis one of your workers. - Extend the Bash hook beyond
cat,headandtail: matchsed -nranges andcat -non big files, let piped and grep commands through. - Run one real question with and without the
.claude/folder, fromclaude -p --output-format json, and comparetotal_cost_usdandduration_ms, not the input column alone. - Check two delegated answers against the source with grep before trusting the reader's summary; budget the main model's verification turn as part of the cost.
- Measure a multi-turn session too: the single-turn results leave the main context at 50k to 119k tokens versus 84k to 414k without delegation, so the second question should start cheaper, but that was not measured.
Go further
- Read the hook format and the
agent_typefield in the official reference before copying a hook from a blog post; the deny shape and the fields on stdin are what make the subagent exemption possible s5. - The subagent documentation covers the
modelfrontmatter field and tool restrictions, which is how a reader can be kept read-only and cheap s6. - Spotify's own "What doesn't work" section is the most useful part of the post: no delegated editing (no reliable line numbers in summaries), no delegated reasoning (a missed thread-safety bug), 10 to 30 second round trips s1.
- The Shunt README shows the three-layer structure (hooks, scripts, skills) and the skill texts that tell the model when to delegate; the skill prose is the part worth adapting, not the hook s2.
- Portal modes are a configuration layer over a model plus a system prompt; the same idea maps onto a Claude Code agent file with a
modelfield s4. - The Hacker News thread is where the measurement questions were first raised, and it is a good checklist of what to ask of any token-saving claim s7.
- A Reddit thread documents an orchestrator spawning five workers on its own expensive model; pin the model in the agent file and, if you want a hard guarantee, deny Agent calls without a
modelfield s8.
Sources
- Portal by Spotify cut my Claude Code token usage by 90%, Spotify Engineering. Why read it: the original claim, the three-layer design and an honest limitations section that undercuts the headline.
- Shunt plugin (spotify/portal-ai-plugins), GitHub. Why read it: the actual hooks, scripts and skill texts, short enough to read in full.
- check-file-size hook source, GitHub. Why read it: the 350-line check in a few lines of shell, the template for your own deny.
- Portal Modes documentation, Spotify. Why read it: what a "mode" is, so you can see why it maps to a subagent file.
- Claude Code hooks reference, Anthropic. Why read it: the current deny format and the stdin fields, including the one that identifies a subagent.
- Claude Code subagents, Anthropic. Why read it: the
modelfrontmatter field and tool allowlists for a cheap read-only worker. - Hacker News discussion of the Spotify post, Hacker News. Why read it: the questions about what the 90% measures, asked before anyone re-measured.
- Fable spawned five Fable agents instead of Opus (r/ClaudeCode), Reddit. Why read it: the failure mode a pinned
modelfield prevents.
FAQ
Is the 90% number wrong?
It measures one thing: estimated input tokens in the main context for bulk-read scenarios on large Java files. On the same kind of metric the rebuild saw 41.9% to 79.4% on read scenarios. It says nothing about cost, time or answer quality, and the post does not claim otherwise.
Do I need Portal to get this?
No. The routing lives in a CLAUDE.md rule, two agent files with a pinned model and a PreToolUse hook. Portal supplies the worker models at Spotify; a model: haiku line does the same job in plain Claude Code.
When does delegation cost more?
When the files are small. The 45-line test-writing scenario cost 2.6% more with delegation because the writer is a second full context and the main model still re-read and tested the result. Every delegated run was also slower, +65.3% on average.
Why did the hook never fire?
Because the routing rule in CLAUDE.md made the main model check wc -l and delegate before it tried to read. The hook only matters when the model ignores the rule, which happened in the verification run without CLAUDE.md.
AIDive