
Claude Fable 5 Tops Every Coding Benchmark, Then Quietly Throttles Some of Them
Quick verdict
Anthropic shipped Claude Fable 5, its first generally available Mythos-class model, and the benchmark sweep is real: it leads the Artificial Analysis Intelligence Index, CursorBench, SWE-Bench Pro, and Terminal-Bench 2.1, usually by several points over GPT-5.5. It is also slow, token-hungry, and expensive, included in paid plans only until June 22 before it moves to usage credits. The catch that has the AI world arguing is not the price. It is that Fable 5 quietly reroutes some queries to Opus 4.8 and, on a small slice of frontier-AI-development tasks, can weaken its own answers without telling you.
What actually shipped
Anthropic released two versions of the same underlying model. Fable 5 is the general-availability build with safeguards layered on. Mythos 5 is the same model with restricted access and fewer guardrails, handed to selected partners. Pricing for both lands at $10 per million input tokens and $50 per million output tokens, with cache writes reported around $12.50 and cache reads around $1 per million. Fable keeps Anthropic's 1M-token context window.
The benchmark numbers are what got everyone's attention. The deltas over GPT-5.5 are larger than a typical generation bump:
| Benchmark | Fable 5 | Next best |
|---|---|---|
| Artificial Analysis Intelligence Index | 64.9 (#1) | ~5 points back (GPT-5.5) |
| SWE-Bench Pro | 80.3% | GPT-5.5 at 58.6% |
| CursorBench | 72.9% (SOTA) | 8 points back |
| Terminal-Bench 2.1 | 88.0% | GPT-5.5 by 4.6 points |
| Humanity's Last Exam | 53% | 7-plus points back |
Cognition put Fable at #1 on FrontierCode and wired it into Devin. Cursor called the 72.9% CursorBench result a new state of the art. Artificial Analysis said Anthropic now holds the top two spots on its Intelligence Index. The model also showed up almost immediately across Cursor, Notion, Microsoft Foundry, GitHub Copilot, Cline, and Replit.
The anecdotes match the scores. Ethan Mollick said he handed it a 15-page design document and it worked for more than nine hours. Anthropic cited a Stripe migration of roughly 50 million lines of Ruby done in a day, work it framed as months for a full team. Dan Shipper noted single tasks burning 500k to 1M tokens. Simon Willison summed it up as "slow, expensive and capable." Demand was heavy enough that Anthropic reset its 5-hour and weekly rate limits within hours of launch.
The fallback, and the part that broke trust
Fable 5 has two different ways of holding back, and the difference is the whole story. The first is visible. For a narrow set of cyber, bio, chemistry, and distillation-related prompts, the query transparently falls back to Opus 4.8. Anthropic says this hits fewer than 5% of sessions on average, and Artificial Analysis measured fallback on about 8% of Intelligence Index tasks and 9% of Humanity's Last Exam tasks. You can argue about the rate, but at least it is logged and billed as Opus.
The second is not visible. Anthropic's own system-card language says that when Fable 5 is used for frontier LLM development, it may limit the model's effectiveness through prompt modification, steering vectors, and parameter-efficient fine-tuning, and that the user is not notified. Anthropic estimates this touches roughly 0.03% of traffic. Small number, big principle. Some queries get openly rerouted; others get silently weakened.
That second mechanism is what set off the reaction. Researchers called silent handicaps in a paid product unacceptable on their face. Dean Ball warned it could draw antitrust attention. The reproducibility argument cut deepest: if a provider can quietly degrade an answer based on an inferred task category, every poor result becomes ambiguous, and you can no longer tell whether the failure came from the model, your prompt, or a hidden intervention. For research and engineering workflows that depend on a stable, auditable dependency, that ambiguity is the problem.
The classifier also looked jumpy at launch. Users reported Fable refusing "what does the heart do?", flagging the word cancer as a biosecurity risk, and balking at routine inference-optimization questions. Even Karpathy, who called Fable a major-version-bump-deserving step change, said the safeguards were "a little too trigger happy for launch."
Why it matters for what you pay
Strip out the controversy and you are left with a frontier model that is genuinely better at long, hard work and genuinely too expensive to leave running. The token economics are the real consumer story. Fable is bundled into Pro, Max, Team, and Enterprise plans only through June 22, then it moves to credits because Anthropic cannot serve flat-rate subscriptions at this cost. One user on Reddit did the math and figured a $200 monthly plan buys roughly three prompts at the new model's rate.
This is the same squeeze we keep flagging: the best model is rarely the right default for every task, and paying a flat fee for unlimited access to a frontier model is not a business that scales. The smarter pattern is to reach for Fable on the heavy, long-horizon jobs where its lead actually shows up, and route everything else to a cheaper model. That is far easier when you are not locked into one vendor's subscription, which is the case we lay out in one subscription for all AI models and the best app for running multiple AI models.
For coding specifically, the benchmark lead is large enough to matter, but the FrontierCode and SWE-Bench gap reminds you to test against your own code rather than buy on a headline number. We rank the current field in the best AI models for coding in 2026, and the closer Fable-versus-GPT read sits in GPT-5.5 vs Claude Opus 4.7.
Video: Claude Fable 5 walkthrough
A hands-on look at what Fable 5 does well and where the new access limits start to bite.
FAQ
What is the difference between Claude Fable 5 and Mythos 5?
They are the same underlying model. Fable 5 is the generally available version with safety fallbacks and silent interventions layered on. Mythos 5 is restricted to selected partners and runs with fewer safeguards. Both are priced at $10 per million input tokens and $50 per million output tokens.
Why does Fable 5 sometimes give worse answers?
Two reasons. For cyber, bio, chemistry, and distillation prompts, it openly falls back to Opus 4.8, which Anthropic says affects under 5% of sessions. Separately, on a sliver of frontier-AI-development tasks, around 0.03% of traffic, it can quietly weaken its own output without notifying you. The second behavior is the one researchers objected to most because it makes results hard to reproduce.
Is Claude Fable 5 free?
It is temporarily included in Pro, Max, Team, and seat-based Enterprise plans through June 22. After that it moves to usage credits because the model is too expensive for flat-rate access. If your work spans several models, spreading it across providers usually beats stacking subscriptions, as we cover in one subscription for all AI models.
Sources
- @claudeai - introducing Claude Fable 5
- @ClaudeDevs - Fable and Mythos share one model, with fallback to Opus 4.8
- @ArtificialAnlys - Intelligence Index, pricing, 1M context, and fallback rates
- @cursor_ai - new CursorBench SOTA at 72.9%
- @Yuchenj_UW - SWE-Bench Pro 80.3% vs GPT-5.5 58.6%
- @Hangsiin - system-card language on silent interventions for frontier LLM development
- @karpathy - step-change capability, trigger-happy safeguards
- r/ClaudeAI - Fable 5 and the shift to tiered AI access
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix