← All posts / Models

The Mid-Tier inversion: Claude Sonnet 5.5 Outscores Opus 5.5 on Agentic Coding While Cutting Task Costs Up to 30%

Anthropic's new mid-tier model scores 70.6% on Terminal-Bench 4.0 — above Opus 5.5 — while generating output 30%+ faster and costing up to 30% less per task. It is also the first Sonnet shipped with cyber safeguards and anti-distillation classifiers.

The Mid-Tier inversion: Claude Sonnet 5.5 Outscores Opus 5.5 on Agentic Coding While Cutting Task Costs Up to 30%

One day before OpenAI’s DevDay, Anthropic quietly reshuffled the economics of its model lineup. On September 28, 2026, the company released Claude Sonnet 5.5, the second model in the Claude 5.5 family — and the headline number is not a benchmark record. It is a mid-tier model outscoring its own flagship sibling on agentic coding: 70.6% on Terminal-Bench 4.0 versus Opus 5.5’s 66.4%, at half the list price.

Sonnet 5.5 lands just a week after Opus 5.5 (September 22) and roughly three months after Sonnet 5, whose selling point was efficient agentic deployment. The new model keeps that thesis and pushes it further: it generates output 30%+ faster than Sonnet 5 — making it the fastest Sonnet model to date — and, by needing far fewer tokens for the same work, costs up to 30% less per task despite unchanged list pricing.

What the benchmarks say

Anthropic’s published numbers show a dramatic jump over Sonnet 5, and near-parity with Opus 5.5 on knowledge work:

  • Terminal-Bench 4.0 (agentic coding): 70.6%, versus 10.3% for Sonnet 5 — a roughly sevenfold improvement — and 66.4% for Opus 5.5 at its best effort level.
  • FrontierCode 1.1 (Main): Anthropic says Sonnet 5.5 at High effort matches GPT-6 Sol’s best score for about a fifth of the cost per task. At Max effort it posts 46.2% — slightly below its own Xhigh result, a quirk the company attributes to FrontierCode’s penalty for out-of-scope changes when the model’s code-review skill splits reviews across subagents.
  • CursorBench 4.0: 55.5%, within about two points of Opus 5.5’s 57.8% and far ahead of Sonnet 5’s 34.1%.
  • GDPval-AA v2.1 (real-world work across 44 occupations): 1844 — just two points below Opus 5.5 (1846) and roughly 400 above Sonnet 5 (1449). GPT-6 Sol sits at 1487.
  • Humanity’s Last Exam (with tools): 64.5% versus Opus 5.5’s 67.7%.
  • OSWorld 2.1 (computer use, partial): 80.1% versus 81.8%.
  • Chartography (visual chart recognition, no tools): 61.6% versus Sonnet 5’s 15.6% and GPT-6 Sol’s 53.6%.

There is also a first that is equal parts meme and milestone: Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots.

TechCrunch’s read on the Opus-beating Terminal-Bench score is that agility matters more than raw intelligence for agentic work — a cheaper model can spawn multiple parallel agents without blowing cost limits, an option that is punitive on a flagship-tier model.

The efficiency story is the price cut

List pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output, $0.20 for cache reads, and $2.50 for cache writes (Opus 5.5 is $4/$20). The real discount is in token consumption. Anthropic’s cost-per-task analysis, plotted across five effort levels (Low, Medium, High, Xhigh, Max), claims:

  • On Terminal-Bench at Medium effort — the default in the Claude apps — Sonnet 5.5 far exceeds Sonnet 5’s best score for less than a tenth of the cost per task.
  • On CursorBench at Low effort, it beats Sonnet 5’s best for under a tenth of the cost.
  • On FrontierCode at High effort, it scores ten points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task.

Early testers corroborate the token diet. Balyasny Asset Management ran a private suite of 2,441 finance tasks and measured about 121k tokens per answer versus 497k for Sonnet 5. Slack saw ~14% fewer output tokens on Slackbot evals, Zendesk processed tickets 20% faster, Box reported 2.4x faster and 12% fewer tokens, Lovable counted a third fewer tool calls and roughly half the shell runs, and Atlassian expects Rovo Agents to run up to 30% faster.

The developer anecdotes skew the same way. Epic Games’ COO Daniel Vogel said the model “cleared the same quality bar you’d expect from a higher-tier model” while holding up on multi-hour system-design audits. SpaceXAI’s director of ML Sualeh Asif reported that across 118 real app builds, Sonnet 5.5 produced apps scoring level with Opus 5 — in 3.6 iterations per build on average, where Opus 5 took 7.7. Unity said it completed 90% of tasks in a multi-step Unity Editor benchmark.

Safety firsts for a mid-tier model

Because Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5’s, Anthropic is shipping it with guardrails previously reserved for flagship models. It is the first Sonnet to launch with cyber safeguards and fallbacks: higher-risk cybersecurity tasks will visibly fall back to Sonnet 5, while routine development is unaffected. An expanded Cyber Verification Program will offer tiered access to advanced capabilities across Sonnet 5.5, Opus 5.5, and Claude Mythos models.

It is also the first Sonnet with safety classifiers that prevent reasoning extraction — a defense against industrial-scale distillation attacks — and it expands “preserved thinking,” binding Claude’s chain of thought to the account that created it (a migration note for developers who move conversations between accounts).

On Anthropic’s automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures. More striking: on containment evaluations, it comes close to Opus 5.5 — the best model tested — in how rarely it attempts sandbox escape, and it is the least likely of any Anthropic model to probe the limits of its container. Given the industry’s current scrutiny of runaway agents, that claim is doing real reputational work.

Why this matters

Two structural shifts are worth watching. First, the mid-tier inversion: when a $2/$10 model beats the flagship on agentic coding benchmarks, “which model is smartest” stops being the buying question — “which model finishes the task cheapest” takes over. Second, the price war has moved from list prices to tokens per task. GPT-6 Sol also lists at $2/$10, so headline pricing between OpenAI and Anthropic mid-tiers is now identical; the competition is purely about how many tokens a task burns. Sonnet 5.5’s fourfold token reduction on real finance work is exactly the kind of number that moves enterprise budgets.

Sonnet 5.5 is available now on the Claude Platform as claude-sonnet-5-5, on AWS, Google Cloud, and Microsoft Azure, with zero data retention. Defaults differ by surface — Medium effort in the Claude apps and Claude Code, High on the platform — and developers running Sonnet with thinking off must switch to the new between_tools setting before migrating. Claude Haiku 5.5, the high-volume tier, follows in the coming weeks.

The message to OpenAI the day before DevDay is hard to miss: the flagship announcements get the attention, but the mid-tier models — fast, cheap, and now safety-hardened — are where the actual workloads live.