Intro: the week that got shorter
Claude Code's weekly limit dropped 17% in mid-September 2026, when the summer promotion ended. People on the biggest plans now report an empty week by Wednesday. Anthropic's own announcement says the limits were permanently raised by 25%, and the post right after it calls the same change a 17% reduction.
Every list of tips for stretching the limit comes without a single number. This article gives each fix a measured one and ranks them. Two results stand out: almost half of a month of tokens went to subagents, and a single long break makes the next message rewrite most of the session.
What changed, and how to count
The measurements here come from one month of Claude Code logs from a single developer: 455 sessions and 63,398 requests, from September 3 to October 3.
The promotion ran from May to September 13 and made the weekly limit 50% higher. The limit on each five-hour window never moved. At the end of August, Anthropic's developer account announced a permanent raise of 25%, and one post later the same thread says it works out to a 17% reduction. Both statements are true:
| Period | Weekly limit (old limit = 100) |
|---|---|
| Before the promotion | 100 |
| During the promotion (May to September 13) | 150 |
| Permanent level since September 14 | 125 |
From 150 down to 125 is the 17% people feel. One user with two of the biggest plans wrote that being at 100% on a Wednesday had never happened before. Another, on the same plan, was at 86% on a Tuesday morning. The cut is not the only cause: a hungrier model shipped at the start of September, so not every empty week comes from this change.
From your side you can see a percentage. The /usage screen splits recent usage between skills, subagents, plugins and each connected MCP server, and it flags cache misses. One key switches it between the last day and the last seven. What you cannot see is the size of the limit in tokens: Anthropic publishes percentages and multipliers, never a token count. Everything measured below is therefore in tokens, from one workload, and not a share of your week.
Counting tokens from the logs has a trap. The log writes the same answer several times, so adding every line gives 18.6 billion tokens. Counted once, it is 9.3 billion. A naive count nearly doubles everything.
Subagents: almost half the bill
A subagent is another Claude that your session starts for a side job, and it reports back when it is done. In the measured month, subagents took 48.1% of all tokens across 2,631 runs.
| Measure | Subagents' share |
|---|---|
| All tokens | 48.1% |
| Output tokens | 63.9% |
| Weighted the way the public price list weighs output and cache writes | 55.3% |
Each subagent also pays an entry price. Before it does anything, its opening request already carries a median of 47,117 tokens: the instructions, the tool list and the skill list, all sent again. Someone else measured it on a different machine and found 16,000 to 21,000 tokens per launch for agents whose own prompt is tiny. As that article puts it, the agent file is a rounding error inside its own launch cost.
The model is the other half. By default a subagent inherits the model of the main conversation, so switching the session to the biggest model puts every helper on it too. In the measured logs, the smallest model handled less than 1% of subagent requests. The fix is one line in the agent's file: a model field set to a smaller model for jobs like running tests or searching files.
That gives two habits. Skip the subagent for a small job you could do in place, and pin a small model on the ones you keep.
The limit of this result: nobody has measured what pinning saves as a share of the week, and a small model that needs more turns can cost more. The 48% comes from work that fans out a lot. Your own share is in the /usage screen.
The five-minute cache nobody lists
Claude Code keeps your conversation in a prompt cache on the server, and reading it back costs a fraction of sending it again. For the main session that cache lives one hour. For a subagent it lives five minutes.
The documentation says it plainly: subagents get five minutes, even on a subscription, until you choose longer. The same goes for everything outside the main conversation, including background work and compaction. The measured logs agree: every cache write from a subagent landed on the five-minute tier, and every one from a main session on the one-hour tier.
A developer on Reddit noticed what that does: one of his subagents rewrote its entire context eight times in a single day. The fix is one line in the settings file, "subagentPromptCacheTtl": "1h".
| His measurement | Before | After |
|---|---|---|
| Cache writes | 12.2 million tokens | 3.0 million tokens |
| Five-hour window with four subagents | from 2% to 100% | from 0% to 22% |
That is one user comparing two different days, not a controlled test. In the logs measured here it barely matters: only 95 of 41,790 subagent follow-up requests (about two in a thousand) came after a wait of more than five minutes, though each one rewrote about 75,000 tokens.
So it depends on how your subagents work. If they wait on a long build, on a review or on you, turn it on. If they run in short bursts, leave it, because a cache that lasts an hour costs more to write.
The break that rewrites the whole session
The main session's cache lasts one hour. After a longer break it is gone, and the next message cannot read anything back. The documentation spells it out: the message you send after the break misses the cache and reprocesses your full context.
| Pause before the message | Requests | Cache rewritten (median) |
|---|---|---|
| Under 5 minutes | 18,029 | 1,176 tokens |
| 5 to 60 minutes | 414 | 1,327 tokens |
| Over 60 minutes | 79 | 130,332 tokens |
The typical session at that moment held 175,523 tokens, so most of it was written again. The meter does not treat a write like a read either. A developer put a logging proxy in front of Claude Code and watched his five-hour window: by his ratios, a token written to the cache weighs about forty times a token read from it.
Claude Code knows this. When you resume a large session after a long break, it offers to resume from a summary instead. Take it.
The cheaper habit comes earlier. When a task is done, clear the session while the cache is still warm. Clearing costs nothing and the next task starts small. Compacting works too, but compacting a huge session is itself a huge request.
A break is not the only way to lose the cache. Switching models in the middle of a session empties it, because each model keeps its own. On the newest models, changing effort does not. Claude Code asks you to confirm a model switch while the cache is warm, and that prompt is the warning.
The limits: 79 cold returns is a small sample, and some of them follow a compaction. A summary also loses detail, so this fix costs some continuity.
Effort: the fix that can cost you quality
Effort is how long the model is allowed to think before it answers. There are five levels, from low to max, and the thinking is billed as output. The default is high on most models and medium on the two newest.
The documentation says the thinking budget can reach tens of thousands of tokens per request and that the top level is prone to overthinking. On the newest models thinking cannot be turned off at all, so the level is the only control.
One developer ran the same 29 real tasks at all five levels:
| Effort level | Mean cost per task | Tasks passed (of 29) |
|---|---|---|
| low | $2.50 | 23 |
| medium | $3.15 | 28 |
| high | $5.01 | 26 |
| xhigh | $6.51 | 25 |
| max | $8.84 | 27 |
Quality did not follow cost. Medium passed more tasks than any level above it, and per dollar it also delivered the most passes. In his words, the curve appears to peak at medium. The Claude Code team works the same way: one of its engineers builds on low or medium, reviews, and only runs verification on high.
The catch is why this fix can cost quality. On the hard problems he picked, low effort passed zero times out of five and high passed five out of five. A low attempt took two minutes, a high one thirty-three.
So match effort to the step: medium to build, high when a mistake is expensive (a bug in old code, a migration, a closing check), max almost never. Session logs record the effort of every request, so you can check what you really ran.
Those costs are in dollars on an older model, not a share of the week; nobody has published that. And a cheap attempt that fails and runs twice costs more than one that works.
The tips that weigh less than advertised
Some fixes are on every list and barely move the needle. They are free to try. They are just not where the week went.
What loads at the start is the real one in this group. One article measured the opening request from an empty folder at 29,061 tokens, and at almost 39,000 inside a real project. In the logs measured here the median opening request is 55,989 tokens, ranging from 15,764 to 105,020 depending on the project. The /context command shows what is in there (memory files, skills, tool lists) and names each memory file it loaded. Trim what you never use. The gain is modest because that block is written once and read from the cache on every turn after that. It hurts on a cold start and on every subagent launch.
| Popular tip | Measured |
|---|---|
| Removing MCP servers | 1,350 tokens for 51 tools across three servers; 18 tokens for one server with one tool |
| Turning off prompt suggestions | 3 to 4% for one user; the "up to 10%" claim came from one account with enormous contexts |
| Filtering shell output | about 0.1% of total volume, measured by a contributor to one of those filters |
Tool definitions are deferred by default now, which is why MCP servers weigh so little. The documentation calls the cost of prompt suggestions small. All three grow with the size of your context, and servers cost more on older models where deferral is off. Switch them off if you like, but do not expect the week back.
The ranked table
Ranked by what was measured:
| Rank | Fix | Measured | Catch |
|---|---|---|---|
| 1 | Fewer, cheaper subagents | 48.1% of tokens; 47,117 per launch | Less parallelism |
| 2 | Don't resume a cold session | 130,332 tokens rewritten against 1,176 | A summary loses detail |
| 3 | Effort: medium to build | $3.15 against $5.01 per task; 28 of 29 passed | Low fails on hard problems |
| 4 | Subagent cache at one hour | 12.2 million to 3.0 million cache-write tokens | Pays only if subagents wait |
| 5 | Trim what loads at the start | +9,744 tokens on a 29,061 baseline | Paid once per session |
The popular three (MCP servers, prompt suggestions, shell output) are not where the week went.
The limits, plainly: this ranking is in tokens, from one month of one person's work, plus other people's measurements. Anthropic does not publish the size of the limit in tokens, so nobody outside can turn these into a share of your week. Your order may differ, and the /usage screen will tell you.
The two biggest fixes are habits, not settings, and they are free: launch fewer subagents, and never resume a cold session in full.
AIDive