AIDive

Video pack

Spotify's 90% Claude Code token cut, rebuilt and measured: hook, subagents, protocol

11 min read

TL;DR

  • Spotify's "90%" is the mean of bulk-read scenarios measured in estimated input tokens on a Java monorepo. The post gives no dollar figure and no quality score.
  • Rebuilt in plain Claude Code (one PreToolUse hook, two cheap subagents, a three-line routing rule) and measured on Fastify over four scenarios and 16 runs, the pattern cut the main model's context by 59.6% and total cost by 33.1%.
  • The deny hooks fired zero times in the measured runs. The savings came from the routing rule in CLAUDE.md; the hooks are the safety net for the day the model ignores it.
  • Delegation was slower every time, +65.3% wall time on average. On the small test-writing scenario it cost 2.6% more.
  • Two traps: hooks also fire inside subagents, so exempt your workers, and sed -n range reads pass straight through a hook that only watches cat, head and tail.
  • The Haiku reader's summary carried factual errors in two of eight delegated runs. Keep the main model's verification turn.

What the measurements say

Spotify's plugin, Shunt, routes bulk work away from the main model through two "modes": a bulk-reader and a code-writer, both running Gemini 2.5 Flash in the examples, with the model field accepting any model configured in the Portal instance s1. The routing has three layers. A check-file-size hook fires on every Read and blocks files over a configurable line threshold, default 350, pointing the model to the bulk-reader skill instead; a check-bash-read hook catches cat, head, tail, less and more on large files while piped commands pass through s1. The hook sources and the two skills are in the public repo s2, with the size check readable on its own s3. The modes themselves live in Portal, Spotify's internal platform, which is why the plugin cannot be run outside the company as shipped s4.

The benchmark claim is thin. Spotify tested four scenarios on a Java monorepo, "measuring tokens Claude would consume reading files directly vs. consuming the bulk-reader's summary", and reports mean bulk-read savings of around 90% s1. The post itself says the code-write scenario is harder to measure in tokens, that worker summaries do not include reliable line numbers so editing cannot be delegated, that the worker missed a subtle thread-safety bug the main model caught, and that each delegation adds 10 to 30 seconds with Portal capping a single invocation at 30 seconds s1. The Hacker News thread raised the same questions about what the 90% measures s7.

The rebuild replaces Portal modes with two Claude Code subagents whose definition files pin the model: an Explore reader on Haiku and a code-writer on Sonnet s6. The deny is a PreToolUse hook returning the deny decision in the current hook JSON format s5. The repo under test was fastify/fastify at commit ac28821d, 294 .js/.ts files, 78270 lines, 63 files over 350 lines. Model ids as reported by the session JSON: main conversation claude-opus-5[1m], reader claude-haiku-4-5-20251001, writer claude-sonnet-5. Each scenario ran twice per config, 16 measured runs, single-turn claude -p sessions with project-only settings so both sides had an identical system prompt s5.

Where the pattern won: S2, a three-file call-graph question over lib/route.js (691 lines), lib/reply.js (1090) and lib/request.js (398), went from a mean main context of 357165.5 tokens to 73440.0 (-79.4%) and from 0.5810500000000001 USD to 0.21823605000000001 USD (-62.4%). Where it did not: S4, writing a test for a 45-line source against a 19-line reference, cost 0.29465575 USD without delegation and 0.3022213 USD with it (+2.6%), because Sonnet is a second full context (13004 to 18729 cache-read tokens) and the main model still re-read the generated file and ran the test s6.

Three findings matter more than the percentages. First, the hooks fired zero times across the 16 measured runs: with the routing rule in CLAUDE.md, the main model ran wc -l and delegated on its own. The only observed deny came from a verification run without CLAUDE.md, where the model was refused Read, then refused cat -n, and answered from grep -n alone without ever calling the Agent tool s5. Second, hooks run inside subagents: in two discarded runs the Haiku reader was itself denied by the size check and fell back to chunked offset/limit reads. The fix is a case "$agent_type" in Explore|code-writer) exit 0 escape at the top of the hook, with the field name confirmed from the logged stdin s5. Third, in the baseline config the main model never used the Read tool at all. It read every file through Bash (cat -n, sed -n '1,200p', sed -n '200,560p'), so a hook that only watches Read catches nothing, and a Bash hook that only matches unpiped cat, head and tail still lets sed -n ranges through s3.

Quality was checked against the source with grep. The baseline produced wrong line numbers in one S2 run (it dumped files with sed -n without line numbers and counted by hand). The delegated config produced three factual errors in the other S2 run and two in an S3 run, all traced to the Haiku summary being taken at face value: wrong callers for buildRequest/buildReply, a non-exported constant listed as an export, a covered iterator marked uncovered. Where the main model spent output tokens re-verifying with grep (S3, 4534 to 4738 output tokens), the answers held s6. On the test-writing scenario every generated file passed: 12/12, 6/6, 7/7 and 10/10 tests. One Reddit report shows the related failure mode of a main model spawning workers on the wrong model when nothing pins it s8, which is what the require-model hook and the model: field in the agent files guard against.

Measurements

Means of the 2 runs per cell. "Main context" is input + cache_creation + cache_read tokens billed to the main model over the session, the figure comparable to Spotify's "tokens in the main context". A = plain Claude Code, B = hook + subagents + CLAUDE.md rule.

scenario main context A main context B change main output A main output B change total cost A total cost B change duration A s duration B s change
S1 88693.0 51551.5 -41.9% 1424.5 1060.5 -25.6% 0.13910675 0.08653685 -37.8% 21.817500000000003 43.799499999999995 +100.8%
S2 357165.5 73440.0 -79.4% 3835.0 2555.5 -33.4% 0.5810500000000001 0.21823605000000001 -62.4% 51.637 129.036 +149.9%
S3 303807.5 114135.5 -62.4% 6192.0 4636.0 -25.1% 0.451037 0.3738534 -17.1% 93.321 124.64099999999999 +33.6%
S4 143431.5 121818.0 -15.1% 5275.5 3340.0 -36.7% 0.29465575 0.3022213 +2.6% 65.7125 86.857 +32.2%
all 4 223274.375 90236.25 -59.6% 4181.75 2898.0 -30.7% 0.366462375 0.2452119 -33.1% 58.122 96.08337499999999 +65.3%

Protocol: two byte-identical shallow clones of fastify/fastify at ac28821d; repo-shunt adds .claude/ (settings, two agent files, three hooks) and a CLAUDE.md routing rule, nothing else. Every session: claude -p "<prompt>" --output-format json --setting-sources project --strict-mcp-config with an empty MCP config, no --model, 600 s timeout. Four prompts, identical on both sides: S1 exports of lib/reply.js, S2 call graph across three lib files, S3 methods of lib/hooks.js versus test/hooks.test.js coverage, S4 write test/head-route.test.js following test/noop-set.test.js. Figures are read from the session JSON modelUsage and total_cost_usd, unrounded. Answers were checked against the source with grep; generated tests were run with node --test.

Do this Monday

  • Run wc -l over your repo and count files over 350 lines. If the count is near zero, stop here: the threshold exists because delegation costs more than it saves on small files.
  • Add a three-line routing rule to your CLAUDE.md: files over the threshold go to a reader subagent, pattern-following code goes to a writer subagent, debugging and architecture stay with the main model. In the measurements this rule did all the work.
  • Create .claude/agents/Explore.md with model: haiku and .claude/agents/code-writer.md with model: sonnet in the frontmatter, so the worker model is pinned in the file and not left to the orchestrator.
  • Write the PreToolUse hook on Read as a safety net, returning the deny decision in the current hook JSON format, and make its first lines exit 0 when agent_type is one of your workers.
  • Extend the Bash hook beyond cat, head and tail: match sed -n ranges and cat -n on big files, let piped and grep commands through.
  • Run one real question with and without the .claude/ folder, from claude -p --output-format json, and compare total_cost_usd and duration_ms, not the input column alone.
  • Check two delegated answers against the source with grep before trusting the reader's summary; budget the main model's verification turn as part of the cost.
  • Measure a multi-turn session too: the single-turn results leave the main context at 50k to 119k tokens versus 84k to 414k without delegation, so the second question should start cheaper, but that was not measured.

Go further

  • Read the hook format and the agent_type field in the official reference before copying a hook from a blog post; the deny shape and the fields on stdin are what make the subagent exemption possible s5.
  • The subagent documentation covers the model frontmatter field and tool restrictions, which is how a reader can be kept read-only and cheap s6.
  • Spotify's own "What doesn't work" section is the most useful part of the post: no delegated editing (no reliable line numbers in summaries), no delegated reasoning (a missed thread-safety bug), 10 to 30 second round trips s1.
  • The Shunt README shows the three-layer structure (hooks, scripts, skills) and the skill texts that tell the model when to delegate; the skill prose is the part worth adapting, not the hook s2.
  • Portal modes are a configuration layer over a model plus a system prompt; the same idea maps onto a Claude Code agent file with a model field s4.
  • The Hacker News thread is where the measurement questions were first raised, and it is a good checklist of what to ask of any token-saving claim s7.
  • A Reddit thread documents an orchestrator spawning five workers on its own expensive model; pin the model in the agent file and, if you want a hard guarantee, deny Agent calls without a model field s8.

Sources

FAQ

Is the 90% number wrong?

It measures one thing: estimated input tokens in the main context for bulk-read scenarios on large Java files. On the same kind of metric the rebuild saw 41.9% to 79.4% on read scenarios. It says nothing about cost, time or answer quality, and the post does not claim otherwise.

Do I need Portal to get this?

No. The routing lives in a CLAUDE.md rule, two agent files with a pinned model and a PreToolUse hook. Portal supplies the worker models at Spotify; a model: haiku line does the same job in plain Claude Code.

When does delegation cost more?

When the files are small. The 45-line test-writing scenario cost 2.6% more with delegation because the writer is a second full context and the main model still re-read and tested the result. Every delegated run was also slower, +65.3% on average.

Why did the hook never fire?

Because the routing rule in CLAUDE.md made the main model check wc -l and delegate before it tried to read. The hook only matters when the model ignores the rule, which happened in the verification run without CLAUDE.md.