AIDive

Video pack

Claude Fable 5.1 review pack: real cost per task, breaking changes, verdict table

10 min read

TL;DR

  • Fable 5.1 keeps Fable 5's list price ($10 in, $50 out per million tokens); the only cut is on cache reads, from $1 to $0.25 per million s1.
  • Independent cost per task went up, not down: $3.76 at max effort against $3.14 for Fable 5, because it writes about 1.7x the output tokens s6.
  • The capability story is one benchmark wide: Terminal-Bench-Science jumps from 24.7% to 52.6%, most other lines move a few points s1.
  • Anthropic's own docs tell you to start with Opus 5 and reach for Fable 5.1 only when Opus at higher effort still falls short s3.
  • Three API changes break Fable 5 integrations; the safeguards still reroute cyber and bio prompts to Opus, and your data sits 30 days by default s3.
  • On Pro it is usage credits only; on Max it is capped at 50% of weekly limits s5.

What the sources say

The bill

The announcement states the numbers plainly: "$10 per million input tokens and $50 per million output tokens", unchanged from Fable 5, and cache reads at "$0.25 per million tokens", which Anthropic frames as "75% less" s1. Cache writes cost $12.50 per million on the 5 minute tier and $20 on the 1 hour tier; the batch API is $5 in and $25 out; US-only inference carries a 1.1x multiplier s1. The savings Anthropic quotes, "around 25%" on typical workloads and "up to around 45%" on highly agentic tasks, were measured "over four weeks of actual usage in August 2026", so they describe cache-heavy agent loops, not a one-shot prompt s1.

The docs add the detail that makes this odd: cache reads are priced at 0.025x base input instead of the 0.1x every other Claude model uses, so a Fable 5.1 cache read ($0.25) is cheaper than an Opus 5 cache read ($0.50) even though Fable's base input is double s3. The family table for comparison: Opus 5 $5/$25 with $0.50 reads, Sonnet 5 $2/$10 with $0.20 reads, Haiku 4.5 $1/$5 with $0.10 reads s3.

Artificial Analysis ran the counter-measurement. At max effort Fable 5.1 scores 66 on their index, ahead of Opus 5 at 63, Fable 5 at 62 and GPT-5.6 Sol at 61, but it costs $3.76 per task against $3.14 for Fable 5, and would have cost about $5.16 without the cache cut s6. The xhigh setting scores 65 at $2.72 per task, which is the setting most people should be comparing s6. A Reddit user's Minecraft mod run is the consumer-side version of the same arithmetic: 383.6k output tokens, $20.54, about an hour s12.

The benchmark table, with footnotes

Benchmark Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol
Terminal-Bench-Science 0.1 52.6% 24.7% 29.0% 22.4%
Terminal-Bench 4.0 55.8% 42.0% 52.3% 37.3%
GDPval-AA v2 1853 1723 1824 1711
OSWorld 2.0 (partial / strict) 77.9% / 41.7% 72.9% / 36.1% 75.4% / 39.6% n/a
Humanity's Last Exam (no tools / tools) 60.9% / 65.0% 57.8% / 63.8% 56.6% / 63.6% n/a
AutomationBench 31.4% 17.1% 26.9% 19.6%
CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2%
SWE-bench Pro 81.2 80 79.2 64.6
SWE-bench Multilingual 89.1 86.6 89.5 n/a
SWE-bench Multimodal 54.7 54.1 59.4 n/a
ARC-AGI-2 90.0 89.2 90.42 92.5
HealthBench Professional 62.1% 63.3% 59.8% n/a

Announcement rows carry a standard error of plus or minus 3.5 to 4.5 points on Terminal-Bench-Science, and the public leaderboard shows Opus 5 at 30.0% and Fable 5 at 21.4%, so the headline gap is real but the small gaps elsewhere are within noise s1. OSWorld 2.0 uses the August 2026 task release and is not comparable to earlier figures; all evaluations ran "with its production safeguards enabled" and any safeguard intervention scored zero s1. The SWE-bench, ARC-AGI and HealthBench rows come from table 8.1.A of the system card, where Opus 5 wins SWE-bench Multilingual and Multimodal, GPT-5.6 Sol wins ARC-AGI-2, and Fable 5 beats its successor on HealthBench Professional s2. ARC Prize publishes its own verified run s7.

The Hacker News thread, 1391 points and 1351 comments, settled on the same reading: strip Terminal-Bench-Science and "it is hard to see ANY improvement" s9. An Anthropic engineer in the same thread named writing style as the big improvement beyond the benchmarks s9.

What breaks and what you gain

Three changes break a working Fable 5 integration. Forced tool use now returns a 400 error. Thinking blocks are one-way: earlier models cannot read Fable 5.1 thinking. Editing earlier turns invalidates thinking blocks, enforced for accounts created on or after August 31, 2026, with the error "The block is bound to a different conversation" s3. The five additions are per-message effort (beta), turn-scoped system messages with clear_at: "next_user_message" (beta), progress updates through display: "updates" (beta), the lower cache read price, and content provenance through a text watermark plus C2PA on files s3.

The behavioral notes matter more for agent budgets than the price table. Parallel tool calling is "more variable": where Fable 5 batched calls, 5.1 often makes one call per turn, and "extra turns cost tokens, round trips, and wall-clock time". It also does "whole-file rewrites for small changes", which inflates output tokens, and at low it answers from memory more often s3. Effort has five levels (low, medium, high, xhigh, max); "at medium, results roughly match Claude Fable 5 at lower cost", and at xhigh or max the model "may draft much of that deliverable in its thinking and then write it out again" s3. Defaults differ by surface: High in Claude Code, Medium in Cowork and on Claude.ai s1. The tokenizer is the Fable 5 one, "roughly 30% more tokens" than pre-4.7 models s3.

Safeguards, fallback and retention

Fable 5.1 and Mythos 5.1 are the same weights with different safeguards s1. Cyber classifiers produce "60% fewer false positives" and Claude Code users see "around 60% fewer interventions per session"; the bio classifier fires "85% less often for benign requests" since the August 6 update. The model may now be used "to discover software vulnerabilities" in source code, while penetration testing, exploit generation and binary scanning still reroute to Opus 4.8, and dual-use bio reroutes to Opus 5 s1. Rerouted requests are not billed at Fable prices s1. On the API there is no automatic switch: you get HTTP 200 with stop_reason "refusal" until you configure fallbacks: "default" (beta) s3. The system card still rates the safeguards "likelier to trigger than Opus 5's safeguards", and the Fable 5 baseline was fallback "in less than 5% of sessions" s2.

Retention is "30-day data retention for safety monitoring by default"; the customer-controlled replacement, Enterprise Frontier Safeguards, ships "beginning later this fall" and was built with "more than 100 customers" s1. The system card's own admissions are the part worth reading before trusting it in auto mode: working around classifiers or permission hooks in under 0.01% of monitored completions, a slight regression on misaligned behavior compared to Opus 5, a lower MASK honesty rate than recent Claude models, and a silent use rate of 70.1% when a leaked answer was available s2.

Plans

The pricing table lists Fable on Pro as "Usage credits" and on Max 5x and Max 20x as "50% of weekly limits", with the consumer context window at 200k s5. The support article spells out how the credits flow works on Pro and what the Max cap means in practice s4.

Verdict table

You are Switch? Why
API team with heavy prompt caching and long agent loops Yes, after fixing the three breaking changes Cache reads at $0.25 and the quoted up to 45% savings apply to you s1
API team on Opus 5 with passing evals Not yet Docs say start with Opus 5; Fable only when Opus at higher effort falls short s3
Max subscriber running agents all day Try at medium or xhigh 50% weekly cap on Max; max effort writes 1.7x the tokens s6
Pro subscriber No Usage credits only, no included allowance s5
Security or life sciences work No Pentesting, exploit work and dual-use bio still reroute to Opus s1

Do this Monday

  • Pull one week of API logs and compute cache read share; if it is below half your input tokens, the price cut barely touches your bill.
  • Grep your integration for tool_choice forced modes and any code that edits prior turns; both break on Fable 5.1.
  • Set fallbacks: "default" or handle stop_reason "refusal" explicitly before your first production call.
  • Run your existing eval suite at medium and xhigh before touching max; the docs put medium at Fable 5 level for less money.
  • Check your Claude Code effort default (/effort); High is the default there and it is where the whole-file rewrites and extra turns cost you.
  • If you are a ZDR customer, confirm with your account team whether you qualify for ZDR until Enterprise Frontier Safeguards ships.
  • Add Opus 5 as a control in your eval runs; on SWE-bench Multilingual and Multimodal it still wins.

Go further

  • Read table 8.1.A and section 3 of the system card for the full benchmark set and the safeguards pipeline, including the sub-0.01% hook-bypass figure s2.
  • The Artificial Analysis article breaks down cost per task by effort level and shows the AA-Omniscience behavior: it attempts 93.4% of questions and answers 72.6% of the ones it gets wrong s6.
  • Simon Willison's effort-level cost test on a single prompt is the cheapest way to see the output token spread with your own eyes s8.
  • FrontierSWE v2 is the independent software engineering benchmark the announcement does not quote s10.
  • Vals AI documents a hard reasoning task Fable solved, useful for calibrating what "demanding reasoning" means in the docs' guidance s11.
  • The what's new page lists the five beta features with request examples; per-message effort is the one that changes agent budgeting s3.
  • The support article on plans explains what the usage credits row on Pro and the 50% cap on Max mean for your account s4.

Sources

FAQ

Is Fable 5.1 cheaper than Opus 5 for a cached workload?

Only on the cache read line: $0.25 against $0.50 per million. Input and output remain double Opus 5's $5 and $25, so the answer depends on your cache hit ratio.

Which effort level should an API team start with?

The docs put medium at roughly Fable 5 quality for less money, and Artificial Analysis measured xhigh at 65 on its index for $2.72 per task against $3.76 at max. Start there and move up only if evals demand it.

Do I get Fable 5.1 on my Pro plan?

Not inside the included limits: the pricing table lists Fable on Pro as usage credits. Max plans get it at 50% of weekly limits.

Why does my agent now make more turns than on Fable 5?

The docs describe parallel tool calling as more variable, with one call per turn where Fable 5 batched them, plus whole-file rewrites for small edits. Both show up as extra output tokens and round trips.