One Model, Two Faces: Inside Anthropic's Claude Fable 5.1 and the Locked-Down Mythos 5.1
Anthropic's Claude Fable 5.1 doubles its science benchmark score, cuts cache-read pricing 75%, and ships with a twin — Mythos 5.1 — that is the same model with weaker safeguards, restricted to vetted cybersecurity and life-sciences organizations.
While the AI world spent the week arguing about OpenAI’s GPT-6 Astra, Anthropic quietly shipped the most structurally interesting model release of the season: Claude Fable 5.1, a general-availability frontier model for demanding reasoning and long-horizon agentic work, and Claude Mythos 5.1, its identical twin — same weights, same capabilities — available only to vetted cybersecurity and life-sciences organizations through restricted-access programs.
The two-models-one-brain design is the story. Anthropic is explicitly acknowledging that its most capable model is too useful in dangerous domains to ship unrestricted, and too dangerous to fully unlock for everyone. Fable 5.1 ships with robust safeguards for cybersecurity and biology queries — many of which are automatically routed to less capable models when classifiers trip. Mythos 5.1 removes most of those guardrails for a small population of trusted defenders and researchers. It is the clearest production-scale experiment yet in capability gating: not “safe model vs. dangerous model,” but the same intelligence wearing two different sets of locks.
The benchmark jump
The headline number is Terminal-Bench Science, where Fable 5.1 scores 52.6%, up from 24.7% for Fable 5 — more than doubling in one generation, and far ahead of Claude Opus 5 at 29.0% and GPT-5.6 Sol at 22.4% (Anthropic reports a standard error of roughly 3.5–4 points). On independent testing, the picture is similarly strong: Artificial Analysis’ Intelligence Index v4.1.1 puts Fable 5.1 at 66 versus 61 for GPT-6 Astra, a five-point lead for the Claude camp — and in Claude Code, Fable 5.1 tops the index at 70.
The comparison with Astra is not entirely one-sided, and honesty requires the asterisks. Astra, released a day later, reclaimed state-of-the-art on several agentic benchmarks — one analysis puts it at 64.6% on Terminal-Bench Science, ahead of Fable 5.1’s 52.6%, and it leads on Terminal-Bench 4.0 and DeepSWE. The two companies are now trading the crown benchmark-by-benchmark rather than model-by-model, which is itself a sign of how compressed the frontier has become.
The safeguard tax
Here is what makes the system card fascinating reading. Fable 5.1 was evaluated with production safeguards active — safety classifiers on, with fallback to Claude Opus 5 for biology-adjacent queries. On tasks where those safeguards intervened, the model scored zero, including on general agentic benchmarks like OSWorld. Across the Artificial Analysis Intelligence Index, fallback routing accounted for roughly 4% of output tokens.
Think about what that means: every published Fable 5.1 benchmark is a net number that already includes the cost of its own safety layer. Anthropic’s model isn’t just competing with Astra’s raw capability — it’s competing with one hand on its own brake, and still winning on several indexes. Karo Zieminski aptly dubbed this the “safeguard tax,” and it is measurable in a way safety overhead has rarely been before. Anthropic’s own support documentation is blunt that users doing routine, legitimate cybersecurity work with Fable 5.1 “should expect high fallback rates” — a frank admission that the classifier net catches dolphins along with sharks.
Pricing: the 75% cache cut
The release also carries an aggressive economic signal. Cache-read pricing drops 75%, and Anthropic estimates Fable 5.1 will cost about 25% less than Fable 5 for typical token-billed workloads — with savings “often much larger” for cache-heavy agentic applications. For the long-horizon agents Fable is explicitly built for, where a single task may re-read the same context hundreds of times, cache economics dominate total cost. The price cut is effectively a subsidy for exactly the workload class Anthropic wants to win: multi-hour, multi-step autonomous work.
The backstory nobody should forget
Fable 5.1’s safeguards are not theoretical architecture — they are scar tissue. The original Fable 5 launched June 9, 2026 as Anthropic’s first “Mythos-class” model deemed safe for public use. Within days it was jailbroken via a multi-agent attack strategy, and on June 12 — three days after launch — Anthropic suspended the model entirely amid U.S. national-security scrutiny and export-control entanglement. It returned July 1 only after the U.S. lifted controls tied to the incident, armed with a new classifier that Anthropic says blocks the original jailbreak method in over 99% of attempts.
That episode reshaped how the industry talks about jailbreaks — Anthropic introduced a three-tier severity classification (minor, narrow harmful, broad) — and it explains why 5.1 exists at all: a rebuilt classifier stack, a restricted twin for the professionals who genuinely need the raw capability, and a system card that leads with cyber capability disclosure. The system card states plainly that Fable 5.1 and Mythos 5.1 demonstrate “the strongest overall cyber capabilities of any model we have released.” When OpenAI’s own Astra documentation says it crossed a critical-cyber safeguard threshold the same week, the dual release reads less like product strategy and more like the industry’s first shared acknowledgment that cyber capability has outgrown the single-model, single-policy paradigm.
Why it matters
Three takeaways worth watching beyond the benchmark horse race:
Capability gating is now a product category. The Fable/Mythos split institutionalizes “know your customer” for frontier AI. Expect every lab with cyber- or bio-capable models to ship a similar restricted tier within the year — the alternative is the Fable 5 outcome: total suspension for everyone.
Safety overhead is now a published metric. The 4% fallback token rate and the zeroed OSWorld scores give external auditors a concrete number to track across generations. If safeguard overhead shrinks while capability grows, that is real safety progress; if it grows, the tax argument against safe deployment gets empirical ammunition.
Cache economics decide the agent era. A 75% cache-read cut targets the exact cost structure of long-running agents. Model selection in 2026 is increasingly determined not by leaderboard deltas of five points, but by what a 12-hour autonomous run costs at the margin.
Fable 5.1 is available now across Claude platforms and APIs wherever usage is billed by token. Mythos 5.1 is not something you can buy — and that, arguably, is the most important feature it has.
Sources
- [1] https://www.anthropic.com/claude-fable-and-mythos-5-1
- [2] https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf
- [3] https://thezvi.substack.com/p/claude-fable-51-and-mythos-51-the
- [4] https://www.latent.space/p/ainews-claude-fablemythos-51-new
- [5] https://www.rdworldonline.com/anthropic-doubles-a-science-benchmark-score-with-fable-5-1-while-openai-says-its-astra-models-crosses-critical-cyber-threshold/
- [6] https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- [7] https://www.anthropic.com/news/redeploying-fable-5