← All posts / Models

Doubling Science Scores, Splitting Safety: Anthropic's Claude Fable 5.1 and Mythos 5.1

Anthropic's Fable 5.1 doubles its Terminal-Bench-Science score to 52.6% while keeping $10/$50 pricing, and its safeguard-free twin Mythos 5.1 ships to vetted cyberdefenders and life scientists through trusted-access programs.

Doubling Science Scores, Splitting Safety: Anthropic's Claude Fable 5.1 and Mythos 5.1

Just days after OpenAI rolled out GPT-6 Astra, Anthropic has answered with a release that is less about a single headline number and more about a structural split in how frontier AI gets deployed. On September 1, 2026, the company introduced Claude Fable 5.1 and Claude Mythos 5.1 — two versions of what is essentially the same underlying model, separated not by capability but by who is allowed to use which capabilities, and under what safeguards.

Fable 5.1 is generally available to all Claude API customers and subscription users. Mythos 5.1, by contrast, ships only through Anthropic’s trusted-access programs to vetted cybersecurity defenders and life scientists in the United States. The pairing turns what could have been a routine point-release into a live experiment in tiered model deployment — and the early benchmark numbers suggest the experiment is working.

The science jump: 24.7% to 52.6%

The headline result is Terminal-Bench-Science 0.1, an agentic benchmark that asks models to carry out multi-step scientific research tasks in a live environment. Fable 5.1 scores 52.6%, more than doubling Fable 5’s 24.7% and pulling well clear of Opus 5’s 29.0% and GPT-5.6 Sol’s 22.4% (with standard error between ±3.5 and ±4.5 points, per Anthropic’s own reporting). Doubling an agentic science score in a single point release is not a normal cadence — it reflects the model’s new adaptive reasoning system, which allocates compute per task rather than applying a fixed reasoning depth to everything.

On Terminal-Bench 4.0, the general agentic coding benchmark, Fable 5.1 scores 55.8% against 42.0% for Fable 5 and 52.3% for Opus 5 — and Mythos 5.1, freed from certain refusal behavior, reaches 60.9% on the same suite. Independent evaluator vals.ai reports that Fable 5.1 ranks #1 on its RSI Index, a five-task benchmark of autonomous LLM research and development. On Artificial Analysis’s Intelligence Index, the Fable 5.1 Adaptive Reasoning Max Effort configuration ranks first out of 192 models currently tracked.

Anthropic’s own framing is characteristically direct: in internal benchmarks, Fable 5.1 “solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition.”

The Mythos split: same model, safeguards removed

The more consequential half of the release is Mythos 5.1. It is the same model as Fable 5.1, but with cyber and life-sciences safeguards deliberately loosened for professional use. According to the system card, Mythos 5.1 “substantially outperforms Claude Opus 5 on almost all cyber evaluations,” including ExploitBench and OSS-related suites. The Zvi review of the system card notes the model is “getting close to” Tier 2 capability — the threshold at which a model can conduct cyber operations completely autonomously, with novel offensive techniques.

That is precisely why access is gated. Mythos 5.1 is available only to a vetted set of US-based cyberdefenders and life scientists through two trusted-access programs. The logic is straightforward: defensive security work and biology research routinely trip ordinary safety filters, because investigating an exploit or a pathogen requires the model to reason about exploits and pathogens. A single safeguarded model forces a choice between refusing legitimate work and enabling misuse. Anthropic’s answer is to split the deployment, not the model.

Pricing: unchanged headline, quietly cheaper

Fable 5.1 keeps Fable 5’s rates: $10 per million input tokens and $50 per million output tokens. The real cost story is caching. Cache reads dropped from $1.00 to $0.25 per million tokens — a 75% cut — which matters enormously for the long-horizon agentic workloads these models increasingly run. When a coding agent holds a million-token context across hundreds of steps, cache pricing dominates the bill. Both models carry a 1,000,000-token context window and support up to 128,000 tokens of output per response.

The context: a crowded week, a skeptical audience

The release landed in one of the busiest weeks in frontier AI memory — OpenAI’s GPT-6 Astra shipped September 3 with explicit AGI-adjacent claims, and Gemini 3.8 Flash arrived at $0.75 per million tokens to attack the price floor. Fable 5.1’s positioning is deliberately different: no AGI rhetoric, no record-shattering claim, just doubled task scores and a governance experiment.

The reception has been mixed in the ways that matter. Enthusiastic API users report the adaptive reasoning modes deliver exactly what long-horizon agents need. Subscription users tell a rougher story — Reddit threads from Max-plan users describe Fable 5.1 as a “resource hog” that burns through usage quotas far faster than Fable 5 did, an artifact of the model choosing higher reasoning effort by default. And a persistent current of benchmark skepticism runs through community discussion, sharpened by the memory of Opus 5’s launch: great numbers, disappointing real-world behavior. “Nobody believes the benchmarks anymore,” one highly upvoted comment put it.

Analysis: the split is the story

Three things make this release worth watching beyond the numbers.

First, tiered deployment is becoming product architecture. Anthropic has effectively turned safety policy into a shipping feature: same weights, two products, different guardrails, gated access. If Mythos 5.1’s trusted-access programs demonstrably accelerate defensive security work without incident, expect every frontier lab to copy the pattern — and expect regulators to start asking who vets the “vetted.”

Second, adaptive reasoning changes the economics of use. Per-task compute allocation means headline token prices matter less than effective cost per completed task. Artificial Analysis already scores Fable 5.1’s Max Effort configuration as offering the lowest cost per task among frontier options at $6.12. The 75% cache-read cut compounds this: the agent workload that defines 2026 is long-context and cache-heavy, and Anthropic just repriced its most important workload.

Third, the science benchmark jump raises the stakes on evaluation. A score doubling on Terminal-Bench-Science 0.1 is either a genuine capability leap or evidence the benchmark saturates quickly — possibly both. With OpenAI simultaneously claiming a “critical cyber threshold” crossed by Astra, the industry is running twin experiments: models getting dramatically more capable at consequential tasks, and benchmarks straining to measure that capability credibly. The RSI Index result — #1 on autonomous R&D — sits exactly at that uncomfortable intersection.

Fable 5.1 is available today across Claude platforms for API and subscription users. Mythos 5.1 remains behind its access wall, a deliberately small deployment for a model whose capabilities Anthropic’s own system card describes as approaching autonomous cyber-operation territory. The gap between those two availability profiles is the real product: a frontier lab betting that the future of powerful AI is not one model for everyone, but the same model, carefully portioned.