TL;DR
- The weekly Claude Code limit moved from a base of 100 to a promo level of 150, then settled at a permanent 125 on September 14, 2026. Against the promo level, that is a 17% cut; against the old base, a 25% raise. Both statements are true at once.
- In one month of local logs, subagents took 48.1% of all tokens and 55.3% of weighted cost. The single biggest lever is launching fewer subagents and pinning a small model on the ones you keep.
- Subagents write a 5-minute cache while the main session writes a 1-hour cache. A follow-up request after a cold gap re-writes about 19 times more cache than a warm one.
- A break longer than 60 minutes in a main session costs a median 130,332 tokens of cache re-write on the next request, against 1,176 when the gap is under 5 minutes.
- Lowering effort did not lower output per request in these logs (mean 778 tokens at high vs 837 at medium on main sessions), so treat it as a quality trade, not a free saving.
- Turning off prompt suggestions and filtering shell output are real but small levers. Count them last.
What the measurements say
The arithmetic behind the headline: base 100, promo level 150, permanent level 125. 125 / 150 = 0.8333, so the cut is 16.67% and rounds to 17%. The wrong reading is to subtract the increments (50% down to 25%) and call it a 25% cut. s2
The promotion ran from May 13, 2026 through September 13, 2026, raised weekly limits by 50% in Claude Code only, and left 5-hour limits untouched. It applied to Pro, Max, Team and seat-based Enterprise plans. s1
The measurements below come from one machine's Claude Code logs: 455 main sessions, 2,631 subagent runs, 63,398 de-duplicated requests between 2026-09-03 and 2026-10-03. The first finding is about counting itself: each request appears on 1.96 log lines on average, so summing every line overstates total tokens by 99.3%. Any script that reads these logs must de-duplicate on (message.id, requestId) first. s11
Subagents are the largest line. De-duplicated, they account for 48.1% of total tokens, 63.9% of output tokens and 55.3% of weighted cost. The first request of a subagent run carries a median prompt of 47,117 tokens before it does anything; the p90 is 52,681 and the max 126,769. Agents with restricted tool sets start much lower (min 5,295). s8
Model choice compounds this. Subagents inherit the main conversation's model unless a model frontmatter, a per-invocation model parameter, or CLAUDE_CODE_SUBAGENT_MODEL says otherwise, and since v2.1.251 the env var alone no longer overrides frontmatter: you need CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1. In the logs, claude-opus-5 alone was 31.5% of all tokens and 35.8% of weighted cost, with 63.2% of that spent inside subagents. s3
Cache tier is decided by where the request runs. In this data, 100.0% of subagent cache writes were 5-minute and 100.0% of main-session cache writes were 1-hour; no request had a mixed split. Inside subagent runs, only 95 of 41,790 follow-up requests (0.2%) arrived after a gap above 5 minutes, but those wrote a mean 74,582 cache_creation tokens against 3,886 for warm requests. The subagentPromptCacheTtl setting and the CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL env var accept 5m or 1h and require Claude Code v2.1.242 or later. s4
An independent run saw the same split: every subagent request written under ephemeral_5m_input_tokens while the parent used ephemeral_1h_input_tokens, and one agent re-writing all 20,971 tokens of its prefix on a request that came after the five-minute window. s6
The main-session equivalent is the long break. Requests arriving under 5 minutes after the previous one wrote a median 1,176 cache_creation tokens (n = 18,029). Between 5 and 60 minutes, 1,327 (n = 414). Above 60 minutes, 130,332 (n = 79), with a median prompt of 175,523 tokens and a p90 of 674,348. The docs confirm that overage billing also drops the main conversation to the five-minute tier. s3
Session start is the fixed cost: the first request of a main session carried a median 55,989 tokens (p90 72,000), with a per-project spread from 15,764 to 105,020 depending on CLAUDE.md and memory size. An earlier public measurement put the floor at about 29k on an empty directory, 30.4k with 3 MCP servers and 38.8k in a real repo. s9
Effort is the lever the docs push and the logs do not reward. On main sessions, high requests produced a mean 778 output tokens against 837 at medium; subagents at high produced 323 against 642. The comparison is confounded (different tasks, models, projects), so it is a reason to be skeptical, not proof. The Claude Code team's own guidance frames effort as where to spend reasoning, not as a budget knob. s10
Two popular tips measured small. Prompt suggestions cost extra requests, and the widely shared "save ~10%" figure is a ceiling, not a typical saving; the setting is promptSuggestionEnabled: false or CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false. s12 Shell output filtering over twenty days cut 66.7 million tokens of output to 24.1 million, but those 66.7 million were 7.4% of the new tokens consumed in the same window. s11
Measurements
Subagent launch cost, first request prompt in tokens, by model:
| Model | n | min | median | p90 | max |
|---|---|---|---|---|---|
| All | 2,631 | 5,295 | 47,117 | 52,681 | 126,769 |
| claude-opus-5 | 1,262 | 36,864 | 43,905 | 48,032 | 50,398 |
| claude-sonnet-5 | 633 | 5,916 | 52,409 | 53,961 | 126,769 |
| claude-opus-5-5 | 426 | 39,408 | 47,189 | 48,362 | 48,883 |
| claude-sonnet-5-5 | 165 | 44,471 | 47,348 | 50,197 | 50,863 |
| claude-fable-5-1 | 102 | 36,551 | 42,593 | 44,286 | 47,385 |
| claude-haiku-4-5 | 42 | 5,295 | 29,636 | 36,714 | 79,190 |
Cache re-write on resume, main sessions, by gap before the request:
| Gap | n | cache_creation median | mean | cache_read median | prompt median |
|---|---|---|---|---|---|
| < 5 min | 18,029 | 1,176 | 2,391 | 184,169 | 186,412 |
| 5 to 60 min | 414 | 1,327 | 5,575 | 221,857 | 225,168 |
| > 60 min | 79 | 130,332 | 241,499 | 25,264 | 175,523 |
Protocol: read every ~/.claude/projects/*/<uuid>.jsonl main session and every */<uuid>/subagents/agent-*.jsonl run, requests from 2026-09-01. De-duplicate assistant lines on (message.id, requestId) and keep one usage record per request. Total tokens = input + output + cache_read + cache_creation; prompt size = input + cache_read + cache_creation. Gap = time from the last log line of the previous request to the first line of this one, inside one session or run. Weighted cost uses relative weights input 1, cache write 5m 1.25, cache write 1h 2, cache read 0.1, output 5; these weights are an assumption, not a published rate.
Do this Monday
- Run
/usageon your plan and read the breakdown by skill, subagent, plugin and MCP, plus the behavior flags raised at 10% or more of recent usage. - List your subagent definitions and add a
model: haikuormodel: sonnetfrontmatter to every one that only searches, checks or summarizes. - If you want one model on every subagent regardless of frontmatter, set
CLAUDE_CODE_SUBAGENT_MODELandCLAUDE_CODE_SUBAGENT_MODEL_FORCE=1. - Check your Claude Code version is 2.1.242 or later, then decide per workflow whether
subagentPromptCacheTtl: "1h"pays: it helps subagents that idle between tool calls, not short ones. - Before a break longer than an hour, finish the task in the current session and write a handoff file; open a fresh session when you come back instead of resuming a 175k-token prompt.
- Write a read-only script over your own logs, de-duplicated on (message.id, requestId), and compare main vs subagent share before changing anything else.
- Set
promptSuggestionEnabled: falseif you never use the suggestions, and treat the saving as a few percent at most.
Go further
- Max 5x against Max 20x: the capacity ratio users measured after the cut is a plan-choice question the video left out. s7
- The full precedence chain for cache TTL (force env, bucket env, bucket setting, subagent
experimental.cacheTtl) and what changes under overage billing. s4 - Why changing effort mid-session can read the whole history with no cache hits on most models, and which models are exempt. s3
- How to read
usagefields and cache tiers in your own session logs, and the ccboard approach to a per-day budget view. s11 - Three subagent files with three models and what each launch actually wrote to cache, with the author's own corrections appended. s8
- Deferred MCP tool definitions and
ENABLE_TOOL_SEARCH=auto:Nto control when tool schemas load into context. s3
Sources
- Claude Code May to August 2026 weekly limits promotion, Anthropic help center. Why read it: the exact dates, plans and scope of the 50% promo, in Anthropic's words.
- Anthropic is cutting Claude Code's current weekly limits by 17 percent, BleepingComputer. Why read it: the 25% raise and the 17% cut side by side, with the quoted announcement.
- Manage costs effectively, Claude Code docs. Why read it: the
/usagebreakdown, subagent model inheritance, and the effort and MCP cache rules. - How Claude Code uses prompt caching, Claude Code docs. Why read it: the only authoritative description of the 5m and 1h tiers and the TTL settings.
- Sub-agents burning your Claude Code 5-hour window? Check the cache TTL, Reddit r/ClaudeAI. Why read it: the thread that surfaced the subagent 5-minute cache, with the setting that flips it.
- Max20x is now just 1.5 times better than Max5x, Reddit r/ClaudeCode. Why read it: a user-side capacity comparison between the two Max tiers after the cut.
- Three Claude Code subagent files, three models in frontmatter, dev.to. Why read it: a reproducible launch-cost experiment with raw usage fields and honest corrections.
- Claude Code token overhead: what 33k actually costs, devaireviews.com. Why read it: a start-up overhead measurement on an empty directory, with MCP servers, and in a real repo.
- Using Claude Code: Spending your effort, X, Claude Code team. Why read it: how the team thinks about effort levels; it contains no token counts, which is the point.
- The real cost of AI, part 9: measure your own usage, florian.bruniaux.com. Why read it: a twenty-day log study that puts shell-output filtering in proportion to total consumption.
- PSA: Turn off Prompt Suggestions, save ~10% of your limits/spend, Reddit r/ClaudeAI. Why read it: the original claim and the thread that narrows the 10% to a ceiling.
FAQ
Is it a 17% cut or a 25% raise?
Both, measured from different baselines. Against the pre-promo base of 100, the permanent level of 125 is a 25% raise. Against the promo level of 150 that users had from May 13 to September 13, 2026, it is a 17% cut.
Should I set subagentPromptCacheTtl to 1h everywhere?
Only if your subagents idle for more than five minutes between requests. In the logs, 0.2% of follow-up requests hit that case, so a blanket 1h tier mostly pays a higher write price for nothing. Measure your own gap distribution first.
Does lowering effort save tokens?
Not visibly in these logs: main-session requests at high produced a mean 778 output tokens against 837 at medium. The data is confounded, so the honest answer is that effort is a quality knob whose saving you have to measure on your own tasks.
Why does my first prompt after lunch cost so much?
The main session writes a 1-hour cache. After a gap above 60 minutes, the next request re-writes the prefix: a median 130,332 cache_creation tokens in the logs, against 1,176 for a warm request. Finish tasks before long breaks and start fresh afterwards.
AIDive