Fable-Class at 40% Off: Anthropic's Claude Opus 5.5 Rewrites the Frontier Pricing Playbook
Anthropic's first release since its 'pacing the frontier' call delivers Fable 5.1-class performance at $4/$20 per million tokens, record behavioral-audit scores, and a 40% cost cut that pressures OpenAI on price.
Ten days after Anthropic publicly called for the industry to slow down, it shipped a model that quietly resets the economics of the frontier. Claude Opus 5.5, announced September 22, 2026, is the first release in Anthropic’s new Claude 5.5 family — and the company’s first model since CEO Dario Amodei’s “pacing the frontier” appeal. The irony is the point: pacing, in Anthropic’s telling, doesn’t mean ceding capability. It means pairing frontier-class performance with tighter evaluation, stronger safeguards, and — unusually for this market — a lower price.
What shipped
Opus 5.5 performs “at the level of Claude Fable 5.1 on most work” while costing 40% less to run than its predecessor Opus 5, according to Anthropic. API pricing lands at $4 per million input tokens and $20 per million output tokens — a 20% cut versus Opus 5. The sharper discount is in cache reads, which dominate the bill for agentic and coding workloads: $0.20 per million tokens, 60% below Opus 5’s $0.50. A Fast mode runs at $8/$40 with up to 2.5x speed in Claude Code and the Claude Platform. The model carries a 1M-token context window with 128,000-token max output, and it generates text more than 30% faster than Opus 5. Claude Sonnet 5.5 and Haiku 5.5 follow “in the coming weeks.”
The benchmark table tells a consistent story. On Terminal-Bench 4.0 at xhigh effort, Opus 5.5 scores 66.4% versus 52.3% for Opus 5, 57.9% for GPT-6 Astra, and 37.3% for GPT-5.6 Sol. On FrontierCode v1.1 (main), its default-effort score of 54.4% edges GPT-6 Astra’s 53.3% — at roughly a fifth of the cost per task. CursorBench 4.0: 57.8%, an 11-point margin over GPT-5.6 Sol for about a third of the cost. Knowledge work shows the same pattern: 1846 Elo on GDPval-AA v2.1 (ahead of Fable 5.1’s 1735 and Opus 5’s 1708), and 67.7% with tools on Humanity’s Last Exam. The exceptions are instructive — GPT-6 Astra retains the lead on Terminal-Bench-Science (64.6% vs 58.7%) and narrowly on Zapier’s AutomationBench (41.4% vs 40.0%).
Efficiency is the headline
The single most quotable data point is the HAProxy test. Anthropic asked both Opus 5.5 and Fable 5.1 to translate HAProxy, the widely deployed C load balancer, into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests — but Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1, at 51% lower cost. An early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5x the tokens. Another completed a 680,000-line code migration in less than a day — work that would have taken an engineering team weeks. In a controlled web-app optimization task, Opus 5.5 cut load times successfully 39 out of 40 attempts, where Opus 5’s smaller fixes also altered app behavior.
Customer quotes pile on the same theme. GitHub’s CPO Mario Rodriguez reported Opus 5.5 solved more terminal tasks than Opus 5 in VS Code “in less than half the steps.” Quantium collapsed a task that took 38 prompts over four days down to 11 prompts in three hours. Optiver matched Opus 5’s quality in half the turns, time, and output tokens. Kiro measured ~40% fewer agent calls and half the tokens on command-line tasks. Viktor, an AI employee living in Slack and Teams, saw costs nearly halve while doubling accuracy on its hardest tasks. In a market where agents burn tokens by the hour, token efficiency is the real pricing story — the per-token cut and the tokens-per-task drop compound into that headline 40%.
Safety, measured this time
The “pacing the frontier” framing gets concrete in the safety section. Opus 5.5 was tested before release by external evaluators including Frontier Design and METR, and it posts the best score to date on Anthropic’s automated behavioral audit — an alignment suite spanning thousands of simulated scenarios. It is, per Anthropic, much less likely than recent models to take hard-to-reverse actions or exceed given boundaries, and more resistant to prompt injection than Opus 5, matching or beating it across coding, tool use, computer use, and web browsing settings. On AI security firm Gray Swan’s benchmark, it ties Fable 5.1 for the lowest prompt-injection success rate of any model tested. Anthropic also says it broadened alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents — with the caveat that limits remain, detailed in the Opus 5.5 System Card.
Because Opus 5.5 is “comparable to Claude Mythos 5.1” in biology and cybersecurity capability, it ships with Fable 5.1-grade safeguards: vetted organizations can apply to the Life Sciences Verification Program for biology research now, with the Cyber Verification Program expanding in coming weeks. For enterprises, Anthropic is also pitching “the most secure coding agent” — an action-screening classifier on every step, an open-source sandbox security teams can audit, and pre-merge code review.
Communication as a feature
A quieter but telling change: Opus 5.5 writes differently. Anthropic’s side-by-samples show Opus 5 producing technically correct but meandering bug explanations, while Opus 5.5 leads with the answer, labels sections, and keeps working sessions readable. Testers’ verdict — “it writes the way I do.” That sounds soft until you consider that verifiability is a safety property: output a human can actually follow is output a human can actually check. Walleye Capital’s team saw the model catch an off-by-one error in their own evaluation instructions and correct for it, flagging that doing so would cost it grader points. No prior model had done that.
The competitive read
Two moves matter here beyond raw scores. First, price: a frontier-class model at $4/$20 with $0.20 cache reads puts direct margin pressure on OpenAI’s GPT-6 family, whose Sol and Astra variants Opus 5.5 repeatedly matches or beats at a fifth to half the cost per task. One wire report noted Opus 5.5 outscored GPT-5.6 Sol on a software development benchmark at roughly one-third the price. Second, narrative: Anthropic spent the last month arguing for international coordination on frontier development — and its first post-appeal release demonstrates that “pacing” is compatible with shipping. Whether regulators and rivals read that as leadership or as having it both ways, the model itself is available now across the Claude Platform, Claude Code, AWS Bedrock, and other providers, with Sonnet 5.5 and Haiku 5.5 expected to complete the family shortly.
Sources
- [1] https://www.anthropic.com/news/claude-opus-5-5
- [2] https://platform.claude.com/docs/en/models/opus-5-5/overview
- [3] https://mashable.com/tech/claude-opus-launch-promises-better-performance-cheaper-price
- [4] https://aws.amazon.com/blogs/machine-learning/claude-opus-5-5-is-now-available-on-aws/
- [5] https://llm-stats.com/blog/research/claude-opus-5-5-launch
- [6] https://www.reddit.com/r/Anthropic/comments/1wnecjb/introducing_claude_opus_55_the_first_model_in_our/