70.6% on Terminal-Bench: Anthropic's Claude Sonnet 5.5 Turns the Workhorse Into a Frontier Contender
Anthropic's Claude Sonnet 5.5 jumps from 10.3% to 70.6% on Terminal-Bench 4.0 at unchanged $2/$10 pricing, runs 30% faster, and becomes the first Sonnet to ship with frontier-grade cyber safeguards and distillation defenses.
Six days after shipping Claude Opus 5.5, Anthropic has released the second member of the Claude 5.5 family: Claude Sonnet 5.5, now available across all platforms including Amazon Web Services, Google Cloud, and Microsoft Azure. On paper it is a mid-tier refresh. In practice, the benchmark deltas are anything but mid-tier — and the safety story attached to it may be the more consequential part of the announcement.
The headline number: 10.3% to 70.6%
The single most striking figure in the release is Terminal-Bench 4.0, an agentic coding evaluation that measures how well a model completes complex, multi-step professional tasks inside a command-line environment. Sonnet 5 scored 10.3%. Sonnet 5.5 scores 70.6% — not only a roughly seven-fold jump over its predecessor, but higher than Opus 5.5’s own 66.4% on the same benchmark.
That inversion — a cheaper, faster model outscoring the flagship on a specific agentic eval — is the pattern to watch. The full comparison table from Anthropic:
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 (Main) | 46.2% (Max) | 42.4% | 54.4% | 49.3% (Xhigh) |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 (knowledge work) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (long-horizon) | 1811 | 1359 | 1822 | 1483 |
| Humanity’s Last Exam (with tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (computer use) | 80.1% partial | 57.0% partial | 81.8% partial | — |
| Chartography (chart recognition) | 61.6% no tools | 15.6% no tools | 64.4% no tools | 53.6% no tools |
On knowledge work, Sonnet 5.5 lands within two points of Opus 5.5 on GDPval-AA (1844 vs 1846) while scoring roughly 400 points above its own predecessor. On Chartography, visual chart recognition without tools, it leaps from 15.6% to 61.6%. And in a detail that will delight a certain kind of benchmark watcher: it is the first Sonnet model to beat Pokémon Red working only from screenshots — a proxy for long-horizon, image-understanding persistence.
Artificial Analysis independently placed Sonnet 5.5 at #2 on its Intelligence Index at max effort, an 18-point gain over Sonnet 5, behind only Opus 5.5.
Same price, less work: the economics of the upgrade
Pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads — half of Opus 5.5’s $4/$20 rates. But the sticker price understates the shift, because Sonnet 5.5 typically needs far fewer tokens to finish the same job. Anthropic reports it costs up to 30% less per task than Sonnet 5, and generates output 30%+ faster, making it the fastest Sonnet model to date.
The effort-dial economics are where it gets sharp. On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score for roughly one-tenth the cost per task. At High effort on FrontierCode, it scores ten points higher than Sonnet 5 at the same setting while costing about one-fifteenth per task. Claude Code and the Claude apps default to Medium effort; the Claude Platform defaults to High.
Enterprise testers supplied concrete numbers. Balyasny Asset Management ran a private suite of 2,441 finance tasks and found Sonnet 5.5 used about 121k tokens per answer where Sonnet 5 used 497k. Box reported it 2.4× faster with 12% fewer total tokens. Lovable measured a third fewer tool calls and roughly half the shell runs per task. Atlassian says Rovo Agents will run up to 30% faster. Slack’s Slackbot evals improved “on almost all” measures with about 14% fewer output tokens, unchanged prompts.
CodeRabbit’s independent review found the model caught 6 of 13 known-bug cases against 4 for Sonnet 5, at 41.2% actionable-comment rate. Epic Games said it held up on a system design audit across “tens of thousands of lines” of gameplay architecture code, and Base44 reported it needed 3.6 iterations per app build where Opus 5 needed 7.7.
First Sonnet to ship with frontier-grade safeguards
The quiet half of the announcement is safety architecture migrating downmarket. Because Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5’s, it becomes the first Sonnet model to launch with cyber safeguards and fallbacks of the kind previously reserved for Anthropic’s most capable models: higher-risk cybersecurity tasks will visibly fall back to Sonnet 5, while routine development is unaffected. An expanded Cyber Verification Program will offer tiered access to advanced capabilities on Sonnet 5.5, Opus 5.5, and the Claude Mythos line.
It is also the first Sonnet to launch with safety classifiers preventing reasoning extraction, targeting distillation attacks that use thousands of fake accounts to strip a model’s capabilities at industrial scale. Preserved-thinking coverage expands too — Claude’s chain of thought can no longer be decoupled from the account that created it, a change developers will notice if they move conversations between accounts mid-session in Claude Code.
On alignment, Anthropic’s automated behavioral audit (roughly 1,850 scenarios) shows Sonnet 5.5 improving on or matching Sonnet 5 on most measures. On newer containment evaluations, it comes close to Opus 5.5 in how rarely it attempts sandbox escape and is the least likely of any Anthropic model to probe its container limits. The company notes, as usual, that no evaluation suite catches everything.
The read
Two takeaways. First, the mid-tier is no longer a compromise tier. When a $2/$10 model posts a 70.6% Terminal-Bench score — beating the flagship on that eval — the practical default for well-scoped agentic coding, bug fixing, and document work has shifted overnight. Anthropic itself positions the pairing: Opus 5.5 for complex work requiring sustained judgment, Sonnet 5.5 for everything else, with the two converging at higher effort settings.
Second, safety features are propagating down the stack at release speed, not years later. Cyber safeguards, distillation defenses, and preserved thinking all debuted on frontier models and reached the workhorse tier within the same generation. With Claude Haiku 5.5 “in the coming weeks” to complete the family, the question for buyers is no longer whether the cheap model is safe enough — it’s whether the expensive model is worth 2× per token for work that no longer needs it.
Availability is immediate across Anthropic’s API, AWS, Google Cloud, and Azure, with zero data retention as with its siblings.
Sources
- [1] https://www.anthropic.com/claude-sonnet-5-5
- [2] https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/
- [3] https://artificialanalysis.ai/articles/claude-sonnet-5-5
- [4] https://venturebeat.com/technology/anthropic-launches-claude-sonnet-5-5-with-30-cost-reduction-per-task-due-to-faster-speeds-and-fewer-tool-calls
- [5] https://www.datacamp.com/blog/claude-sonnet-5-5