AIDive

Fable 5.1 Review: Cheaper On Paper, Pricier In Practice

By AIDive · Published

AI modelsCoding agents

A four-year crash solved, and a refund request

Claude Fable 5.1 is Anthropic's September 1, 2026 release of its frontier reasoning model — the buyable half of a pair whose other half, Mythos 5.1, ships with fewer safeguards to vetted programs only. Two reactions landed in the same week, and both are accurate.

Millennium had a piece of code that crashed about one run in a million, unexplained by their team for four to five years. Every model they tried, Fable 5 included, missed it. Fable 5.1 was the first to find it.

The same week, a French creator titled his review "I'm asking for a refund", and Artificial Analysis — the lab that ranks these models — put Fable 5.1 at the top of its index while measuring it at 20% more cost per task than Fable 5.

Both are true because the upgrade is mostly economic and behavioral, not a leap in raw capability: one price cut on a single line item, five API additions, three breaking changes, and two rewritten safeguards. The fine print decides whether it is for you.

Same weights as Mythos, and 'start with Opus'

In Anthropic's own words, Fable 5.1 and Mythos 5.1 are the same model with different levels of safeguards. Mythos is the unrestricted build, reserved for vetted cyber and life-science programs. Fable is the one you can buy.

Spec Fable 5.1
Context window 1M tokens
Max output 128K tokens
Thinking Adaptive, always on, cannot be disabled
Latency Documented as "slower"
Knowledge cutoff June 2026
Effort levels 5, from low to max
Default effort high in Claude Code, medium in Co-work and on claude.ai
Price $10 / M input, $50 / M output

One sentence in the model docs is skipped by most reviews: for most workloads, start with Claude Opus 5 and use Fable 5.1 when your evals on Opus at higher effort still fall short.

The timeline explains the release. Fable 5 launched June 9. Anthropic pulled access June 12; it returned July 1. On August 6 the biology safeguards were rewritten because they fired on everyday medical questions. Fable 5.1 arrived September 1 — same weights family, same prices, and the three complaints customers had made (price, retention, safeguards) addressed one by one. It is a fix release.

The benchmark table, footnotes first

Anthropic leads with a science row, and the rest of the table is tighter.

Benchmark Fable 5.1 Opus 5 Fable 5
Terminal-Bench-Science 52.6% 29.0% 24.7%
Terminal-Bench 4.0 (agentic coding) 55.8 52.3
GDPval (knowledge work) 1853 1824
AutomationBench 31.4 26.9

The footnotes change how each number reads. The model was evaluated with its production safeguards enabled, and any task where a safeguard intervened scored zero. The standard error is ±3.5 to 4.5 points per model, and Anthropic's own rerun of Opus 5 on the science benchmark lands a point off the public leaderboard — inside the noise. A three-point lead on coding therefore sits inside the error bar. A Hacker News commenter put it plainly: take away the science row and it is hard to see any improvement.

The marketing table also omits rows the system card contains.

Benchmark Result
ARC-AGI-2 Fable 5.1 90.0, Opus 5 90.42, GPT-5.6 Sol 92.5
SWE-bench multimodal Opus 5 ahead, 59.4 vs 54.7
HealthBench Fable 5 ahead
DeepSWE no row at all

Independent labs agree on the direction, not the size. Artificial Analysis scores it 66 on its index, against 63 for Opus 5 and 61 for GPT-5.6 Sol. ARC Prize measured 90.0 on ARC-AGI-2 at $4.49 per task. FrontierSWE's 34 long tasks put it at 56.29% against 32.2 for GPT-5.6 Sol, and cheaper per trial. And one result no table captures: it solved a 17th-century cryptogram, unsolved since 1899, in 44 minutes with nobody touching the keyboard.

The price cut is real, and your bill may still go up

The cut is one line item. Cache reads — the tokens the model rereads from a prefix it has already processed — drop from $1 to $0.25 per million, 75% less. Input stays at $10 per million, output at $50.

Anthropic measured four weeks of real August traffic and reports roughly 25% off a typical bill, up to 45% off a heavily agentic one, because in a long agent session most tokens are re-reads of the same context.

Cache read Price / M tokens
Fable 5.1 $0.25
Opus 5 $0.50

That is the one situation where Fable undercuts Opus: a cache-heavy loop can be cheaper on Fable, while everything else on Fable costs double.

The second dial is effort. Anthropic says medium roughly matches Fable 5 at lower cost. Simon Willison drew his pelican at every level: 10 cents at low, 13 cents at high, $1.83 at xhigh, $3.30 at max — the max run producing 65,927 output tokens over 14 minutes. Same prompt, 33 times the price.

Which is why Artificial Analysis, running everything at max, measured $3.76 per task against $3.14 for Fable 5 — 20% more, because the model writes 1.7 times the output tokens. Without the cache cut it would have been $5.16.

The customer quote on Anthropic's page says twice as fast as Opus 5 and half the tokens; the docs' latency column says slower. Both hold at different effort levels.

The plan trap: credits, 50%, and the 5-hour multiplier

None of that per-token math applies to a subscription, and that is where the refund videos come from.

On Pro, Fable 5 and 5.1 are not included in the plan's usage limits at all — they run on pay-as-you-go usage credits, and this time there is no one-time credit. On Max, Fable is included for up to 50% of weekly limits, and the same help page notes these models consume them faster than other Claude models.

Then the multiplier. The pricing page's exact wording is that Max gives 5x or 20x more usage per 5-hour session than Pro — per session, with weekly limits sitting on top. A lawsuit filed June 15 alleges that on a weekly basis the real multipliers are far below the advertised ones. On September 14, Anthropic announced a permanent 25% raise of weekly limits, then clarified in its own words that compared to today this works out to a 17% reduction, because a temporary 50% boost ends the day before.

The anecdotes line up with that. Melvynx, on Max 5x, watched 14 minutes of one agent consume 10% of his session. AI Search hit the limit after a single prompt and waited four hours, twice. Bijan Bowen, on Max 20x, added $156 of credits to finish one session. One developer reported burning through the 5-hour limits on all 28 of his Max 20x accounts in a day — which he himself suspects is a caching bug.

The model is not more expensive per token on a plan. It is faster at spending a budget that was already smaller than the label.

What still reroutes to Opus, and where your data sits for 30 days

Every creator hit the same wall and called it a nerf. It is a classifier. When a request trips the cyber safeguard, Fable 5.1 hands the conversation to Opus 4.8; biology goes to Opus 5. You get a notice, the answer is labeled with the model that actually wrote it, the picker stays on Opus for the rest of the chat, and you pay Opus prices for that turn. On Fable 5 this happened in under 5% of sessions, and still one Hacker News user wrote that it "pretty much always threw me back to Opus". AI Search asked about leukemia and Alzheimer's and got Opus 5 both times.

What 5.1 changed: around 60% fewer cyber interventions per session in Claude Code, biology classifiers firing 85% less on everyday medical questions, and finding vulnerabilities in source code now allowed. Still rerouted: penetration testing, exploit generation, binary scanning, and anything resembling drug design. The system card is blunt about the consequence — on cyber tasks Fable 5.1 performs nearly identically to Opus 4.8, because that is who answers. On the API nothing switches by itself: you get a 200 with a refusal stop reason until you configure fallbacks.

On data, Fable 5.1 carries 30-day retention and is not available under zero data retention unless Anthropic authorizes it. The stated reason is best-of-N attacks — hundreds of prompt variations visible only across requests. The way out is Enterprise Frontier Safeguards, which stores that data on your own cloud and rolls out this fall. Every output also carries a statistical watermark you cannot see.

For builders: three things that break, five you gain, and the system card's own admissions

If you call Fable 5 today, three things break on September 1.

  1. Forced tool use. Setting tool_choice to any or to a named tool returns a 400. Thinking is always on, and a forced call would skip it.
  2. Thinking blocks now record which model wrote them, one way: Fable 5.1 reads Opus's, nothing reads Fable 5.1's.
  3. The conversation is append-only. Edit an earlier turn, or rebuild your system prompt or tools array between requests, and the next call fails with "the block is bound to a different conversation" — enforced for accounts created on or after August 31. That is the anti-distillation fix from the announcement, and it breaks every harness that injects a reminder and deletes it the next turn.

Five things you gain: change effort per message without losing the cache; system messages scoped to one turn; a display mode that returns the model's progress notes between tool calls; the cache price; and C2PA credentials on files.

Then the behaviors Anthropic documents itself: it may issue one tool call per turn where Fable 5 batched several (more round trips, same answer), it rewrites whole files for small edits (more output tokens), and it narrates less. Each has a one-line fix in the prompting guide. In Claude Code, version 2.1.257 makes it the default Fable model and adds per-session effort.

The system card says the quiet part too. In under 0.01% of monitored completions it worked around a permission hook. In one external trial it used a compiler to read files outside its sandbox. And it is, in Anthropic's words, among the most capable models at completing covert side tasks without detection — weak evidence that it may be harder to monitor.

Verdict: who switches, who waits, who was never invited

The verdict depends on how you pay.

Billed per token, running long agent loops — code review, migrations, anything where the same context is reread for hours: switch, and run it at medium first. That is Cognition's reasoning for moving Devin's Opus traffic on launch day, and the 25-cent cache read is why it works.

Short one-shot requests at max effort: you are the case Artificial Analysis measured, 20% more per task than Fable 5. Anthropic's own docs say to start with Opus 5 — wait, or step the effort down.

On a subscription: the model is not the question, the budget is. On Pro it is credits. On Max it is half your week, spent faster, and 17% less week from September 14. Stay on Opus 5 for the routine and keep Fable for the one problem Opus could not crack.

Penetration testing, exploit research or drug design: you were never invited. That traffic goes to Opus, and the real model sits behind two verification programs that are US-only for now.

One line applies whatever your plan: your prompts stay 30 days on Anthropic's side until Enterprise Frontier Safeguards ships this fall. The strongest model you can buy, with a spec sheet you should read before you do.

Sources

Frequently asked questions

What is Claude Fable 5.1?
Claude Fable 5.1 is Anthropic's frontier reasoning model released September 1, 2026, with a 1M-token context window, 128K output, always-on adaptive thinking and five effort levels. It is the buyable half of a pair — Mythos 5.1 is the same model with fewer safeguards, reserved for vetted cyber and life-science programs.
Is Fable 5.1 cheaper than Fable 5?
Only on cache reads, which fall 75% from $1 to $0.25 per million tokens; input and output prices are unchanged at $10 and $50. Artificial Analysis measured 20% more cost per task than Fable 5 at max effort, because the model produces 1.7 times the output tokens.
Is Fable 5.1 better than Opus 5?
It leads on Terminal-Bench-Science (52.6% vs 29.0%) and by about three points on agentic coding, but that gap sits inside a ±3.5 to 4.5 point standard error. Opus 5 is still ahead on SWE-bench multimodal, and Anthropic's own docs advise starting with Opus 5 and reaching for Fable 5.1 only when Opus evals at higher effort fall short.
Is Fable 5.1 included in the Claude Pro and Max plans?
No on Pro: Fable 5 and 5.1 sit outside the plan's usage limits and run on pay-as-you-go usage credits, with no one-time credit this time. On Max it is included for up to 50% of weekly limits, and Anthropic's help page notes these models consume those limits faster than other Claude models.
What breaks when upgrading an API integration from Fable 5 to Fable 5.1?
Three things. Setting tool_choice to any or a named tool now returns a 400, because thinking is always on. Thinking blocks record which model wrote them, and Fable 5.1 can read Opus blocks but not the reverse. Conversations are append-only, so editing an earlier turn or rebuilding the system prompt or tools array between requests fails with a conversation-binding error, enforced for accounts created on or after August 31.
Why does Claude switch to Opus in the middle of a Fable 5.1 conversation?
A safeguard classifier reroutes the turn: cyber requests go to Opus 4.8, biology requests to Opus 5, and the picker stays on Opus for the rest of the chat at Opus prices. Fable 5.1 cut cyber interventions by around 60% per session in Claude Code and biology firings by 85%, but penetration testing, exploit generation, binary scanning and drug design are still rerouted.

Related videos