TL;DR
- Fable 5.1 में Fable 5 की list price वही है ($10 input, $50 output प्रति million tokens); एकमात्र कटौती cache reads पर है, $1 से $0.25 प्रति million s1।
- स्वतंत्र माप में प्रति task cost घटी नहीं, बढ़ी है: max effort पर $3.76, जबकि Fable 5 पर $3.14, क्योंकि यह लगभग 1.7x output tokens लिखता है s6।
- क्षमता की कहानी सिर्फ एक benchmark चौड़ी है: Terminal-Bench-Science 24.7% से 52.6% पर पहुँचा, बाकी ज़्यादातर लाइनें कुछ ही अंक हिलीं s1।
- Anthropic के अपने docs कहते हैं कि Opus 5 से शुरू करें और Fable 5.1 तभी लें जब ऊँचे effort पर भी Opus कम पड़े s3।
- तीन API बदलाव Fable 5 integrations को तोड़ते हैं; safeguards अब भी cyber और bio prompts को Opus पर भेजते हैं, और आपका data डिफ़ॉल्ट रूप से 30 दिन रखा जाता है s3।
- Pro पर यह सिर्फ usage credits से मिलता है; Max पर साप्ताहिक limits के 50% तक सीमित है s5।
सूत्र क्या कहते हैं
बिल
घोषणा में आँकड़े साफ़ लिखे हैं: "$10 per million input tokens and $50 per million output tokens", Fable 5 जैसे ही, और cache reads "$0.25 per million tokens", जिसे Anthropic "75% less" बताता है s1। Cache writes 5 मिनट वाले tier पर $12.50 प्रति million और 1 घंटे वाले पर $20 हैं; batch API $5 input और $25 output है; सिर्फ़ US में inference पर 1.1x multiplier लगता है s1। Anthropic जो बचत बताता है, सामान्य workloads पर "around 25%" और बहुत agentic tasks पर "up to around 45%", वह "over four weeks of actual usage in August 2026" में नापी गई थी, इसलिए यह cache-heavy agent loops का आँकड़ा है, किसी एक बार के prompt का नहीं s1।
Docs एक अजीब बात जोड़ते हैं: cache reads की कीमत base input का 0.025x है, बाकी हर Claude model के 0.1x की जगह, इसलिए Fable 5.1 का cache read ($0.25) Opus 5 के cache read ($0.50) से सस्ता है, जबकि Fable का base input दोगुना है s3। तुलना के लिए family table: Opus 5 $5/$25 और reads $0.50, Sonnet 5 $2/$10 और reads $0.20, Haiku 4.5 $1/$5 और reads $0.10 s3।
Artificial Analysis ने उलटा माप चलाया। Max effort पर Fable 5.1 उनके index पर 66 लाता है, Opus 5 (63), Fable 5 (62) और GPT-5.6 Sol (61) से आगे, लेकिन इसकी cost प्रति task $3.76 है जबकि Fable 5 की $3.14, और cache कटौती के बिना लगभग $5.16 होती s6। xhigh setting 65 लाती है और प्रति task $2.72 लेती है, और ज़्यादातर लोगों को इसी setting से तुलना करनी चाहिए s6। एक Reddit user का Minecraft mod run इसी हिसाब का consumer संस्करण है: 383.6k output tokens, $20.54, लगभग एक घंटा s12।
Benchmark table, footnotes के साथ
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 | 55.8% | 42.0% | 52.3% | 37.3% |
| GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| OSWorld 2.0 (partial / strict) | 77.9% / 41.7% | 72.9% / 36.1% | 75.4% / 39.6% | n/a |
| Humanity's Last Exam (बिना tools / tools के साथ) | 60.9% / 65.0% | 57.8% / 63.8% | 56.6% / 63.6% | n/a |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
| SWE-bench Pro | 81.2 | 80 | 79.2 | 64.6 |
| SWE-bench Multilingual | 89.1 | 86.6 | 89.5 | n/a |
| SWE-bench Multimodal | 54.7 | 54.1 | 59.4 | n/a |
| ARC-AGI-2 | 90.0 | 89.2 | 90.42 | 92.5 |
| HealthBench Professional | 62.1% | 63.3% | 59.8% | n/a |
घोषणा की पंक्तियों में Terminal-Bench-Science पर standard error प्लस-माइनस 3.5 से 4.5 अंक है, और public leaderboard में Opus 5 30.0% और Fable 5 21.4% पर है, इसलिए headline अंतर असली है लेकिन बाकी जगह के छोटे अंतर noise के भीतर हैं s1। OSWorld 2.0 अगस्त 2026 की task release इस्तेमाल करता है और पुराने आँकड़ों से तुलनीय नहीं है; सभी evaluations "with its production safeguards enabled" चले और safeguard के किसी भी हस्तक्षेप को शून्य अंक मिला s1। SWE-bench, ARC-AGI और HealthBench की पंक्तियाँ system card की table 8.1.A से हैं, जहाँ Opus 5 SWE-bench Multilingual और Multimodal जीतता है, GPT-5.6 Sol ARC-AGI-2 जीतता है, और Fable 5 अपने उत्तराधिकारी को HealthBench Professional पर हराता है s2। ARC Prize अपना verified run खुद प्रकाशित करता है s7।
Hacker News thread, 1391 points और 1351 comments, इसी निष्कर्ष पर पहुँचा: Terminal-Bench-Science हटा दें तो "it is hard to see ANY improvement" s9। उसी thread में Anthropic के एक engineer ने benchmarks से परे सबसे बड़ा सुधार writing style को बताया s9।
क्या टूटता है और क्या मिलता है
तीन बदलाव चालू Fable 5 integration को तोड़ते हैं। Forced tool use अब 400 error लौटाता है। Thinking blocks एकतरफ़ा हैं: पुराने models Fable 5.1 की thinking नहीं पढ़ सकते। पिछले turns को edit करने से thinking blocks invalid हो जाते हैं, यह 31 अगस्त 2026 को या उसके बाद बने accounts पर लागू है, error है "The block is bound to a different conversation" s3। पाँच नई चीज़ें ये हैं: per-message effort (beta), clear_at: "next_user_message" वाले turn-scoped system messages (beta), display: "updates" से progress updates (beta), कम cache read कीमत, और text watermark तथा files पर C2PA से content provenance s3।
Agent budgets के लिए व्यवहार वाले नोट्स कीमत की table से ज़्यादा मायने रखते हैं। Parallel tool calling "more variable" है: जहाँ Fable 5 calls को batch करता था, 5.1 अक्सर हर turn में एक call करता है, और "extra turns cost tokens, round trips, and wall-clock time"। यह "whole-file rewrites for small changes" भी करता है, जिससे output tokens फूलते हैं, और low पर यह अधिक बार memory से जवाब देता है s3। Effort के पाँच स्तर हैं (low, medium, high, xhigh, max); "at medium, results roughly match Claude Fable 5 at lower cost", और xhigh या max पर model "may draft much of that deliverable in its thinking and then write it out again" s3। Defaults surface के हिसाब से अलग हैं: Claude Code में High, Cowork और Claude.ai में Medium s1। Tokenizer Fable 5 वाला ही है, 4.7 से पहले के models की तुलना में "roughly 30% more tokens" s3।
Safeguards, fallback और retention
Fable 5.1 और Mythos 5.1 वही weights हैं, बस safeguards अलग हैं s1। Cyber classifiers "60% fewer false positives" देते हैं और Claude Code users को "around 60% fewer interventions per session" दिखते हैं; 6 अगस्त के update के बाद से bio classifier benign requests पर "85% less often" चलता है। Model को अब source code में "to discover software vulnerabilities" इस्तेमाल किया जा सकता है, जबकि penetration testing, exploit generation और binary scanning अब भी Opus 4.8 पर जाते हैं, और dual-use bio Opus 5 पर s1। Reroute हुई requests का बिल Fable की कीमत पर नहीं बनता s1। API पर कोई automatic switch नहीं है: जब तक आप fallbacks: "default" (beta) configure नहीं करते, आपको HTTP 200 के साथ stop_reason "refusal" मिलता है s3। System card अब भी safeguards को "likelier to trigger than Opus 5's safeguards" कहता है, और Fable 5 baseline में fallback "in less than 5% of sessions" हुआ था s2।
Retention "30-day data retention for safety monitoring by default" है; ग्राहक-नियंत्रित विकल्प, Enterprise Frontier Safeguards, "beginning later this fall" आ रहा है और "more than 100 customers" के साथ बना है s1। System card की अपनी स्वीकारोक्तियाँ वह हिस्सा हैं जिसे auto mode में भरोसा करने से पहले पढ़ना चाहिए: निगरानी में रखे गए 0.01% से कम completions में classifiers या permission hooks को दरकिनार करना, Opus 5 की तुलना में misaligned व्यवहार पर थोड़ी गिरावट, हाल के Claude models से कम MASK honesty rate, और leaked जवाब उपलब्ध होने पर 70.1% silent use rate s2।
Plans
Pricing table Pro पर Fable को "Usage credits" और Max 5x तथा Max 20x पर "50% of weekly limits" दिखाती है, consumer context window 200k के साथ s5। Support article बताता है कि Pro पर credits का flow कैसे चलता है और Max के cap का व्यवहार में क्या मतलब है s4।
Verdict table
| आप हैं | Switch करें? | क्यों |
|---|---|---|
| भारी prompt caching और लंबे agent loops वाली API team | हाँ, तीन breaking changes ठीक करने के बाद | $0.25 cache reads और बताई गई 45% तक की बचत आप पर लागू होती है s1 |
| Opus 5 पर चल रही API team, जिसकी evals पास हैं | अभी नहीं | Docs: Opus 5 से शुरू करें; Fable तभी जब ऊँचे effort पर Opus कम पड़े s3 |
| Max subscriber जो दिन भर agents चलाता है | medium या xhigh पर आज़माएँ | Max पर 50% weekly cap; max effort 1.7x tokens लिखता है s6 |
| Pro subscriber | नहीं | सिर्फ़ usage credits, कोई शामिल allowance नहीं s5 |
| Security या life sciences का काम | नहीं | Pentesting, exploit work और dual-use bio अब भी Opus पर जाते हैं s1 |
सोमवार को यह करें
- एक हफ़्ते के API logs निकालें और cache read का हिस्सा गिनें; अगर यह आपके input tokens के आधे से कम है, तो कीमत की कटौती आपके बिल को मुश्किल से छूती है।
- अपने integration में
tool_choiceके forced modes और पिछले turns को edit करने वाले कोड को grep करें; दोनों Fable 5.1 पर टूटते हैं। - पहली production call से पहले
fallbacks: "default"सेट करें या stop_reason "refusal" को साफ़ तौर पर handle करें। -
maxको छूने से पहले अपना मौजूदा eval suitemediumऔरxhighपर चलाएँ; docs medium को कम पैसे में Fable 5 के स्तर का बताते हैं। - अपना Claude Code effort default देखें (
/effort); वहाँ High डिफ़ॉल्ट है और पूरी file की rewrites तथा अतिरिक्त turns का खर्च वहीं पड़ता है। - अगर आप ZDR ग्राहक हैं, तो अपनी account team से पूछें कि Enterprise Frontier Safeguards आने तक क्या आप ZDR के पात्र हैं।
- अपने eval runs में Opus 5 को control के रूप में जोड़ें; SWE-bench Multilingual और Multimodal पर यह अब भी जीतता है।
आगे पढ़ें
- पूरे benchmark set और safeguards pipeline के लिए system card की table 8.1.A और section 3 पढ़ें, जिसमें 0.01% से कम वाला hook-bypass आँकड़ा भी है s2।
- Artificial Analysis का article effort स्तर के हिसाब से प्रति task cost बाँटता है और AA-Omniscience व्यवहार दिखाता है: यह 93.4% सवाल attempt करता है और जो गलत करता है उनमें से 72.6% का जवाब देता है s6।
- Simon Willison का एक prompt पर effort-level cost test output tokens का फैलाव अपनी आँखों से देखने का सबसे सस्ता तरीका है s8।
- FrontierSWE v2 वह स्वतंत्र software engineering benchmark है जिसे घोषणा quote नहीं करती s10।
- Vals AI एक कठिन reasoning task दर्ज करता है जिसे Fable ने हल किया, जो docs की सलाह में "demanding reasoning" का मतलब परखने के काम आता है s11।
- What's new पेज पाँच beta features को request examples के साथ सूचीबद्ध करता है; per-message effort वही है जो agent budgeting बदलता है s3।
- Plans पर support article समझाता है कि Pro पर usage credits वाली पंक्ति और Max पर 50% cap का आपके account के लिए क्या मतलब है s4।
स्रोत
- Claude Fable 5.1 and Mythos 5.1, Anthropic. क्यों पढ़ें: कीमत की table, benchmark table और बचत के footnotes एक ही पेज पर।
- Claude Fable 5.1 and Mythos 5.1 system card, Anthropic. क्यों पढ़ें: अतिरिक्त benchmark पंक्तियाँ और alignment की गिरावट सिर्फ़ यहीं लिखी हैं।
- What's new in Fable 5.1 (model docs), Anthropic. क्यों पढ़ें: breaking changes, effort स्तर और व्यवहार के बदलाव जो आपको कोड में मिलेंगे।
- Claude Fable models on your plan, Anthropic Support. क्यों पढ़ें: Pro, Max और Team को असल में क्या मिलता है।
- Anthropic pricing, Anthropic. क्यों पढ़ें: plan की पंक्तियाँ और 200k consumer context window।
- Claude Fable 5.1 independent analysis, Artificial Analysis. क्यों पढ़ें: हर effort स्तर पर प्रति task cost, quote की हुई नहीं, नापी हुई।
- Anthropic Claude Fable 5.1 on ARC-AGI-2, ARC Prize. क्यों पढ़ें: उस इकलौते benchmark का verified तीसरे पक्ष का run जिसे GPT-5.6 Sol जीतता है।
- Claude Fable 5.1 (pelican effort-level cost test), Simon Willison. क्यों पढ़ें: एक prompt पर effort की तुलना जिसे आप मिनटों में दोहरा सकते हैं।
- Claude Fable 5.1 and Claude Mythos 5.1 (Hacker News discussion), Hacker News. क्यों पढ़ें: token budgets, fallback की आवृत्ति और breaking changes के असली मक़सद पर practitioners की राय।
- FrontierSWE v2, FrontierSWE. क्यों पढ़ें: Anthropic की table से बाहर का एक SWE benchmark।
- Fable solves Cyphral Distich, Vals AI. क्यों पढ़ें: कीमत के सामने तौलने के लिए कठिन reasoning की एक ठोस जीत।
- Fable 5.1 made a Minecraft mod for $20 (r/ClaudeAI), Reddit r/ClaudeAI. क्यों पढ़ें: एक agentic घंटे के लिए असली token और dollar गिनती।
FAQ
क्या cached workload के लिए Fable 5.1 Opus 5 से सस्ता है?
सिर्फ़ cache read की लाइन पर: $0.25 बनाम $0.50 प्रति million। Input और output Opus 5 के $5 और $25 से दोगुने ही रहते हैं, इसलिए जवाब आपके cache hit ratio पर निर्भर है।
API team को किस effort स्तर से शुरू करना चाहिए?
Docs medium को कम पैसे में लगभग Fable 5 की गुणवत्ता पर रखते हैं, और Artificial Analysis ने xhigh को अपने index पर 65 और $2.72 प्रति task नापा, जबकि max पर $3.76। वहीं से शुरू करें और evals माँगें तभी ऊपर जाएँ।
क्या मुझे अपने Pro plan पर Fable 5.1 मिलता है?
शामिल limits के भीतर नहीं: pricing table Pro पर Fable को usage credits के रूप में दिखाती है। Max plans में यह साप्ताहिक limits के 50% पर मिलता है।
मेरा agent अब Fable 5 से ज़्यादा turns क्यों लेता है?
Docs parallel tool calling को अधिक variable बताते हैं, जहाँ Fable 5 calls को batch करता था वहाँ अब हर turn में एक call, साथ में छोटे edits के लिए पूरी file की rewrites। दोनों अतिरिक्त output tokens और round trips के रूप में दिखते हैं।
AIDive