AIDive

Claude Code's Limit Dropped 17%. 5 Fixes, Measured

By AIDive · Published · Updated

Coding agents

Intro: the week that got shorter

Claude Code's weekly limit dropped 17% in mid-September 2026, when the summer promotion ended. People on the biggest plans now report an empty week by Wednesday. Anthropic's own announcement says the limits were permanently raised by 25%, and the post right after it calls the same change a 17% reduction.

Every list of tips for stretching the limit comes without a single number. This article gives each fix a measured one and ranks them. Two results stand out: almost half of a month of tokens went to subagents, and a single long break makes the next message rewrite most of the session.

What changed, and how to count

The measurements here come from one month of Claude Code logs from a single developer: 455 sessions and 63,398 requests, from September 3 to October 3.

The promotion ran from May to September 13 and made the weekly limit 50% higher. The limit on each five-hour window never moved. At the end of August, Anthropic's developer account announced a permanent raise of 25%, and one post later the same thread says it works out to a 17% reduction. Both statements are true:

Period Weekly limit (old limit = 100)
Before the promotion 100
During the promotion (May to September 13) 150
Permanent level since September 14 125

From 150 down to 125 is the 17% people feel. One user with two of the biggest plans wrote that being at 100% on a Wednesday had never happened before. Another, on the same plan, was at 86% on a Tuesday morning. The cut is not the only cause: a hungrier model shipped at the start of September, so not every empty week comes from this change.

From your side you can see a percentage. The /usage screen splits recent usage between skills, subagents, plugins and each connected MCP server, and it flags cache misses. One key switches it between the last day and the last seven. What you cannot see is the size of the limit in tokens: Anthropic publishes percentages and multipliers, never a token count. Everything measured below is therefore in tokens, from one workload, and not a share of your week.

Counting tokens from the logs has a trap. The log writes the same answer several times, so adding every line gives 18.6 billion tokens. Counted once, it is 9.3 billion. A naive count nearly doubles everything.

Subagents: almost half the bill

A subagent is another Claude that your session starts for a side job, and it reports back when it is done. In the measured month, subagents took 48.1% of all tokens across 2,631 runs.

Measure Subagents' share
All tokens 48.1%
Output tokens 63.9%
Weighted the way the public price list weighs output and cache writes 55.3%

Each subagent also pays an entry price. Before it does anything, its opening request already carries a median of 47,117 tokens: the instructions, the tool list and the skill list, all sent again. Someone else measured it on a different machine and found 16,000 to 21,000 tokens per launch for agents whose own prompt is tiny. As that article puts it, the agent file is a rounding error inside its own launch cost.

The model is the other half. By default a subagent inherits the model of the main conversation, so switching the session to the biggest model puts every helper on it too. In the measured logs, the smallest model handled less than 1% of subagent requests. The fix is one line in the agent's file: a model field set to a smaller model for jobs like running tests or searching files.

That gives two habits. Skip the subagent for a small job you could do in place, and pin a small model on the ones you keep.

The limit of this result: nobody has measured what pinning saves as a share of the week, and a small model that needs more turns can cost more. The 48% comes from work that fans out a lot. Your own share is in the /usage screen.

The five-minute cache nobody lists

Claude Code keeps your conversation in a prompt cache on the server, and reading it back costs a fraction of sending it again. For the main session that cache lives one hour. For a subagent it lives five minutes.

The documentation says it plainly: subagents get five minutes, even on a subscription, until you choose longer. The same goes for everything outside the main conversation, including background work and compaction. The measured logs agree: every cache write from a subagent landed on the five-minute tier, and every one from a main session on the one-hour tier.

A developer on Reddit noticed what that does: one of his subagents rewrote its entire context eight times in a single day. The fix is one line in the settings file, "subagentPromptCacheTtl": "1h".

His measurement Before After
Cache writes 12.2 million tokens 3.0 million tokens
Five-hour window with four subagents from 2% to 100% from 0% to 22%

That is one user comparing two different days, not a controlled test. In the logs measured here it barely matters: only 95 of 41,790 subagent follow-up requests (about two in a thousand) came after a wait of more than five minutes, though each one rewrote about 75,000 tokens.

So it depends on how your subagents work. If they wait on a long build, on a review or on you, turn it on. If they run in short bursts, leave it, because a cache that lasts an hour costs more to write.

The break that rewrites the whole session

The main session's cache lasts one hour. After a longer break it is gone, and the next message cannot read anything back. The documentation spells it out: the message you send after the break misses the cache and reprocesses your full context.

Pause before the message Requests Cache rewritten (median)
Under 5 minutes 18,029 1,176 tokens
5 to 60 minutes 414 1,327 tokens
Over 60 minutes 79 130,332 tokens

The typical session at that moment held 175,523 tokens, so most of it was written again. The meter does not treat a write like a read either. A developer put a logging proxy in front of Claude Code and watched his five-hour window: by his ratios, a token written to the cache weighs about forty times a token read from it.

Claude Code knows this. When you resume a large session after a long break, it offers to resume from a summary instead. Take it.

The cheaper habit comes earlier. When a task is done, clear the session while the cache is still warm. Clearing costs nothing and the next task starts small. Compacting works too, but compacting a huge session is itself a huge request.

A break is not the only way to lose the cache. Switching models in the middle of a session empties it, because each model keeps its own. On the newest models, changing effort does not. Claude Code asks you to confirm a model switch while the cache is warm, and that prompt is the warning.

The limits: 79 cold returns is a small sample, and some of them follow a compaction. A summary also loses detail, so this fix costs some continuity.

Effort: the fix that can cost you quality

Effort is how long the model is allowed to think before it answers. There are five levels, from low to max, and the thinking is billed as output. The default is high on most models and medium on the two newest.

The documentation says the thinking budget can reach tens of thousands of tokens per request and that the top level is prone to overthinking. On the newest models thinking cannot be turned off at all, so the level is the only control.

One developer ran the same 29 real tasks at all five levels:

Effort level Mean cost per task Tasks passed (of 29)
low $2.50 23
medium $3.15 28
high $5.01 26
xhigh $6.51 25
max $8.84 27

Quality did not follow cost. Medium passed more tasks than any level above it, and per dollar it also delivered the most passes. In his words, the curve appears to peak at medium. The Claude Code team works the same way: one of its engineers builds on low or medium, reviews, and only runs verification on high.

The catch is why this fix can cost quality. On the hard problems he picked, low effort passed zero times out of five and high passed five out of five. A low attempt took two minutes, a high one thirty-three.

So match effort to the step: medium to build, high when a mistake is expensive (a bug in old code, a migration, a closing check), max almost never. Session logs record the effort of every request, so you can check what you really ran.

Those costs are in dollars on an older model, not a share of the week; nobody has published that. And a cheap attempt that fails and runs twice costs more than one that works.

The tips that weigh less than advertised

Some fixes are on every list and barely move the needle. They are free to try. They are just not where the week went.

What loads at the start is the real one in this group. One article measured the opening request from an empty folder at 29,061 tokens, and at almost 39,000 inside a real project. In the logs measured here the median opening request is 55,989 tokens, ranging from 15,764 to 105,020 depending on the project. The /context command shows what is in there (memory files, skills, tool lists) and names each memory file it loaded. Trim what you never use. The gain is modest because that block is written once and read from the cache on every turn after that. It hurts on a cold start and on every subagent launch.

Popular tip Measured
Removing MCP servers 1,350 tokens for 51 tools across three servers; 18 tokens for one server with one tool
Turning off prompt suggestions 3 to 4% for one user; the "up to 10%" claim came from one account with enormous contexts
Filtering shell output about 0.1% of total volume, measured by a contributor to one of those filters

Tool definitions are deferred by default now, which is why MCP servers weigh so little. The documentation calls the cost of prompt suggestions small. All three grow with the size of your context, and servers cost more on older models where deferral is off. Switch them off if you like, but do not expect the week back.

The ranked table

Ranked by what was measured:

Rank Fix Measured Catch
1 Fewer, cheaper subagents 48.1% of tokens; 47,117 per launch Less parallelism
2 Don't resume a cold session 130,332 tokens rewritten against 1,176 A summary loses detail
3 Effort: medium to build $3.15 against $5.01 per task; 28 of 29 passed Low fails on hard problems
4 Subagent cache at one hour 12.2 million to 3.0 million cache-write tokens Pays only if subagents wait
5 Trim what loads at the start +9,744 tokens on a 29,061 baseline Paid once per session

The popular three (MCP servers, prompt suggestions, shell output) are not where the week went.

The limits, plainly: this ranking is in tokens, from one month of one person's work, plus other people's measurements. Anthropic does not publish the size of the limit in tokens, so nobody outside can turn these into a share of your week. Your order may differ, and the /usage screen will tell you.

The two biggest fixes are habits, not settings, and they are free: launch fewer subagents, and never resume a cold session in full.

Free resources

Sources

Frequently asked questions

Why did Claude Code's weekly limit drop in September 2026?
A promotion that ran from May to September 13 had raised the weekly limit by 50%. When it ended, Anthropic set the permanent level 25% above the old limit, which is 17% below the promotional level people had been using. The five-hour window limit did not change.
What uses the most Claude Code usage?
In one month of measured logs, subagents took 48.1% of all tokens and 63.9% of output tokens. Each subagent launch also carried a median of 47,117 tokens of instructions and tool lists before doing any work. Your own split is shown in the /usage screen.
How do I make a Claude Code subagent use a cheaper model?
Add a model field to the agent's file, for example model: haiku. By default a subagent inherits the model of the main conversation, so a session on the biggest model runs every helper on it too.
What happens when I resume a Claude Code session after a long break?
The main session's prompt cache lasts one hour. After that, the next message misses the cache and reprocesses the full context: in measured logs it rewrote a median of 130,332 tokens, against 1,176 when sent within five minutes. Claude Code offers to resume from a summary instead, which avoids the rewrite.
Which Claude Code effort level should I use?
Medium to build, high when a mistake is expensive, max almost never. In one developer's test of 29 real tasks, medium passed 28 at $3.15 per task while max passed 27 at $8.84. On hard problems, low effort passed 0 of 5 and high passed 5 of 5.
Does removing MCP servers save Claude Code usage?
Very little. Tool definitions are deferred by default, so only names load at the start: one measurement counted 1,350 tokens for 51 tools across three servers, and 18 tokens for a single server with one tool.

Related videos