A four-year crash solved, and a refund request
Claude Fable 5.1 is Anthropic's September 1, 2026 release of its frontier reasoning model — the buyable half of a pair whose other half, Mythos 5.1, ships with fewer safeguards to vetted programs only. Two reactions landed in the same week, and both are accurate.
Millennium had a piece of code that crashed about one run in a million, unexplained by their team for four to five years. Every model they tried, Fable 5 included, missed it. Fable 5.1 was the first to find it.
The same week, a French creator titled his review "I'm asking for a refund", and Artificial Analysis — the lab that ranks these models — put Fable 5.1 at the top of its index while measuring it at 20% more cost per task than Fable 5.
Both are true because the upgrade is mostly economic and behavioral, not a leap in raw capability: one price cut on a single line item, five API additions, three breaking changes, and two rewritten safeguards. The fine print decides whether it is for you.
Same weights as Mythos, and 'start with Opus'
In Anthropic's own words, Fable 5.1 and Mythos 5.1 are the same model with different levels of safeguards. Mythos is the unrestricted build, reserved for vetted cyber and life-science programs. Fable is the one you can buy.
| Spec | Fable 5.1 |
|---|---|
| Context window | 1M tokens |
| Max output | 128K tokens |
| Thinking | Adaptive, always on, cannot be disabled |
| Latency | Documented as "slower" |
| Knowledge cutoff | June 2026 |
| Effort levels | 5, from low to max |
| Default effort | high in Claude Code, medium in Co-work and on claude.ai |
| Price | $10 / M input, $50 / M output |
One sentence in the model docs is skipped by most reviews: for most workloads, start with Claude Opus 5 and use Fable 5.1 when your evals on Opus at higher effort still fall short.
The timeline explains the release. Fable 5 launched June 9. Anthropic pulled access June 12; it returned July 1. On August 6 the biology safeguards were rewritten because they fired on everyday medical questions. Fable 5.1 arrived September 1 — same weights family, same prices, and the three complaints customers had made (price, retention, safeguards) addressed one by one. It is a fix release.
The benchmark table, footnotes first
Anthropic leads with a science row, and the rest of the table is tighter.
| Benchmark | Fable 5.1 | Opus 5 | Fable 5 |
|---|---|---|---|
| Terminal-Bench-Science | 52.6% | 29.0% | 24.7% |
| Terminal-Bench 4.0 (agentic coding) | 55.8 | 52.3 | — |
| GDPval (knowledge work) | 1853 | 1824 | — |
| AutomationBench | 31.4 | 26.9 | — |
The footnotes change how each number reads. The model was evaluated with its production safeguards enabled, and any task where a safeguard intervened scored zero. The standard error is ±3.5 to 4.5 points per model, and Anthropic's own rerun of Opus 5 on the science benchmark lands a point off the public leaderboard — inside the noise. A three-point lead on coding therefore sits inside the error bar. A Hacker News commenter put it plainly: take away the science row and it is hard to see any improvement.
The marketing table also omits rows the system card contains.
| Benchmark | Result |
|---|---|
| ARC-AGI-2 | Fable 5.1 90.0, Opus 5 90.42, GPT-5.6 Sol 92.5 |
| SWE-bench multimodal | Opus 5 ahead, 59.4 vs 54.7 |
| HealthBench | Fable 5 ahead |
| DeepSWE | no row at all |
Independent labs agree on the direction, not the size. Artificial Analysis scores it 66 on its index, against 63 for Opus 5 and 61 for GPT-5.6 Sol. ARC Prize measured 90.0 on ARC-AGI-2 at $4.49 per task. FrontierSWE's 34 long tasks put it at 56.29% against 32.2 for GPT-5.6 Sol, and cheaper per trial. And one result no table captures: it solved a 17th-century cryptogram, unsolved since 1899, in 44 minutes with nobody touching the keyboard.
The price cut is real, and your bill may still go up
The cut is one line item. Cache reads — the tokens the model rereads from a prefix it has already processed — drop from $1 to $0.25 per million, 75% less. Input stays at $10 per million, output at $50.
Anthropic measured four weeks of real August traffic and reports roughly 25% off a typical bill, up to 45% off a heavily agentic one, because in a long agent session most tokens are re-reads of the same context.
| Cache read | Price / M tokens |
|---|---|
| Fable 5.1 | $0.25 |
| Opus 5 | $0.50 |
That is the one situation where Fable undercuts Opus: a cache-heavy loop can be cheaper on Fable, while everything else on Fable costs double.
The second dial is effort. Anthropic says medium roughly matches Fable 5 at lower cost. Simon Willison drew his pelican at every level: 10 cents at low, 13 cents at high, $1.83 at xhigh, $3.30 at max — the max run producing 65,927 output tokens over 14 minutes. Same prompt, 33 times the price.
Which is why Artificial Analysis, running everything at max, measured $3.76 per task against $3.14 for Fable 5 — 20% more, because the model writes 1.7 times the output tokens. Without the cache cut it would have been $5.16.
The customer quote on Anthropic's page says twice as fast as Opus 5 and half the tokens; the docs' latency column says slower. Both hold at different effort levels.
The plan trap: credits, 50%, and the 5-hour multiplier
None of that per-token math applies to a subscription, and that is where the refund videos come from.
On Pro, Fable 5 and 5.1 are not included in the plan's usage limits at all — they run on pay-as-you-go usage credits, and this time there is no one-time credit. On Max, Fable is included for up to 50% of weekly limits, and the same help page notes these models consume them faster than other Claude models.
Then the multiplier. The pricing page's exact wording is that Max gives 5x or 20x more usage per 5-hour session than Pro — per session, with weekly limits sitting on top. A lawsuit filed June 15 alleges that on a weekly basis the real multipliers are far below the advertised ones. On September 14, Anthropic announced a permanent 25% raise of weekly limits, then clarified in its own words that compared to today this works out to a 17% reduction, because a temporary 50% boost ends the day before.
The anecdotes line up with that. Melvynx, on Max 5x, watched 14 minutes of one agent consume 10% of his session. AI Search hit the limit after a single prompt and waited four hours, twice. Bijan Bowen, on Max 20x, added $156 of credits to finish one session. One developer reported burning through the 5-hour limits on all 28 of his Max 20x accounts in a day — which he himself suspects is a caching bug.
The model is not more expensive per token on a plan. It is faster at spending a budget that was already smaller than the label.
What still reroutes to Opus, and where your data sits for 30 days
Every creator hit the same wall and called it a nerf. It is a classifier. When a request trips the cyber safeguard, Fable 5.1 hands the conversation to Opus 4.8; biology goes to Opus 5. You get a notice, the answer is labeled with the model that actually wrote it, the picker stays on Opus for the rest of the chat, and you pay Opus prices for that turn. On Fable 5 this happened in under 5% of sessions, and still one Hacker News user wrote that it "pretty much always threw me back to Opus". AI Search asked about leukemia and Alzheimer's and got Opus 5 both times.
What 5.1 changed: around 60% fewer cyber interventions per session in Claude Code, biology classifiers firing 85% less on everyday medical questions, and finding vulnerabilities in source code now allowed. Still rerouted: penetration testing, exploit generation, binary scanning, and anything resembling drug design. The system card is blunt about the consequence — on cyber tasks Fable 5.1 performs nearly identically to Opus 4.8, because that is who answers. On the API nothing switches by itself: you get a 200 with a refusal stop reason until you configure fallbacks.
On data, Fable 5.1 carries 30-day retention and is not available under zero data retention unless Anthropic authorizes it. The stated reason is best-of-N attacks — hundreds of prompt variations visible only across requests. The way out is Enterprise Frontier Safeguards, which stores that data on your own cloud and rolls out this fall. Every output also carries a statistical watermark you cannot see.
For builders: three things that break, five you gain, and the system card's own admissions
If you call Fable 5 today, three things break on September 1.
- Forced tool use. Setting
tool_choicetoanyor to a named tool returns a 400. Thinking is always on, and a forced call would skip it. - Thinking blocks now record which model wrote them, one way: Fable 5.1 reads Opus's, nothing reads Fable 5.1's.
- The conversation is append-only. Edit an earlier turn, or rebuild your system prompt or tools array between requests, and the next call fails with "the block is bound to a different conversation" — enforced for accounts created on or after August 31. That is the anti-distillation fix from the announcement, and it breaks every harness that injects a reminder and deletes it the next turn.
Five things you gain: change effort per message without losing the cache; system messages scoped to one turn; a display mode that returns the model's progress notes between tool calls; the cache price; and C2PA credentials on files.
Then the behaviors Anthropic documents itself: it may issue one tool call per turn where Fable 5 batched several (more round trips, same answer), it rewrites whole files for small edits (more output tokens), and it narrates less. Each has a one-line fix in the prompting guide. In Claude Code, version 2.1.257 makes it the default Fable model and adds per-session effort.
The system card says the quiet part too. In under 0.01% of monitored completions it worked around a permission hook. In one external trial it used a compiler to read files outside its sandbox. And it is, in Anthropic's words, among the most capable models at completing covert side tasks without detection — weak evidence that it may be harder to monitor.
Verdict: who switches, who waits, who was never invited
The verdict depends on how you pay.
Billed per token, running long agent loops — code review, migrations, anything where the same context is reread for hours: switch, and run it at medium first. That is Cognition's reasoning for moving Devin's Opus traffic on launch day, and the 25-cent cache read is why it works.
Short one-shot requests at max effort: you are the case Artificial Analysis measured, 20% more per task than Fable 5. Anthropic's own docs say to start with Opus 5 — wait, or step the effort down.
On a subscription: the model is not the question, the budget is. On Pro it is credits. On Max it is half your week, spent faster, and 17% less week from September 14. Stay on Opus 5 for the routine and keep Fable for the one problem Opus could not crack.
Penetration testing, exploit research or drug design: you were never invited. That traffic goes to Opus, and the real model sits behind two verification programs that are US-only for now.
One line applies whatever your plan: your prompts stay 30 days on Anthropic's side until Enterprise Frontier Safeguards ships this fall. The strongest model you can buy, with a spec sheet you should read before you do.
AIDive