← All posts / Models

Ninety Percent Off: Claude Haiku 5.5 Completes the 5.5 Family and Resets the Small-Model Price Floor

Anthropic ships Claude Haiku 5.5 at up to 90% below Haiku 4.5 pricing, with first-ever effort controls for the Haiku class, huge agentic benchmark jumps, and a system card that openly discloses safety regressions.

Ninety Percent Off: Claude Haiku 5.5 Completes the 5.5 Family and Resets the Small-Model Price Floor

Six weeks after Claude Opus 5.5 opened a new generation and less than two weeks after Claude Sonnet 5.5 followed, Anthropic has closed the loop: Claude Haiku 5.5 went live on October 7, 2026, the third and final model of the Claude 5.5 family. The small model of the lineup is usually the afterthought of a launch season. This one is the headline, because the pricing attached to it redefines what “cheap” means at the frontier-adjacent end of the market.

The price cut that changes the math

The core numbers are simple. On the Claude Platform, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that threshold, pricing shifts to $0.50 input and $2.50 output per million. Cache reads are $0.01 per million in the low tier; cache writes run $0.125.

Compare that to Haiku 4.5, which launched a year earlier at $1.00 input and $5.00 output. That is a 90% reduction on the sub-100K tier and 50% above it — and Anthropic says roughly 90% of Haiku 4.5 requests fell into that lower tier, which is how the company lands on an “around 75% cheaper on average” figure for typical workloads. For orientation: Sonnet 5.5 still lists at $2.00/$10.00 per million tokens, so short-prompt Haiku 5.5 calls are now a full order of magnitude below the mid-tier family price.

There is an asterisk worth reading. The footnote in Anthropic’s announcement concedes that Haiku 5.5 uses a new tokenizer — similar to the ones in Sonnet 5.5 and Opus 5.5 — and that the model consumes slightly more tokens for the same work. The percentage savings are computed against list pricing, not against final invoice token counts. For most volume workloads the discount survives the tokenizer adjustment comfortably, but anyone doing precise cost modeling should re-baseline their token counts rather than trust a straight ratio.

Positioning: the subagent workhorse

Anthropic is not pretending Haiku 5.5 is a general-purpose brain. The company’s own placement is explicit: high-volume, latency-sensitive tasks — classification, summarization, extraction, routing, compaction — plus real-time experiences like chat, voice agents, and live support. In the 5.5 family division of labor, Opus 5.5 remains the daily driver for complex coding and knowledge work, Sonnet 5.5 covers well-scoped tasks, and Haiku 5.5 is the fast, cheap subordinate that makes multi-agent architectures economically viable.

That framing matters because of where the industry is heading. The dominant pattern of 2026 is orchestration: a planning model that decomposes a task, and a fleet of subagents that execute the pieces. In that world, the unit economics of the small model determine whether the whole architecture pencils out. Rogo’s Applied AI team put it concretely: while a bigger model builds the deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs — accurate enough to trust, cheap enough to run constantly.

Two capability notes stand out. Haiku 5.5 is the first Haiku-class model with adjustable effort controls, letting teams tune cost against intelligence per task — a knob that previously only existed on the larger models. And Anthropic claims it is the company’s fastest model to date at each model’s standard speed, with the caveat that Opus models in Fast Mode still outrun it. Alongside the launch, the Claude Python and TypeScript SDKs gained beta support for computer use and browser use, with Anthropic pointing at Haiku 5.5 as the natural fit for those repetitive automation tasks.

Benchmarks: small model, outsized jumps

The published numbers show the largest generational leap of the three 5.5 launches, precisely because Haiku 4.5 set a low bar on agentic work:

  • GDPval-AA v2.1 (knowledge work): 1620 vs 735 for Haiku 4.5 — and above GPT-6 Luna’s 1437
  • AA-Briefcase v1.1: 1578 vs 614
  • OSWorld 2.1 offline subset (computer use): 72.4% vs 15.7%
  • Humanity’s Last Exam: 45.9% without tools, 57.4% with tools, vs 10.2%/18.7%
  • Terminal-Bench 4.0 (agentic coding): 39.2% vs 0.0% — Haiku 4.5 literally scored zero
  • FrontierCode 1.1 (Main): 46.4%, and 46.4% on Chartography without tools vs 6.4%

Early customer data reinforces the story. AlphaSense measured a statistically significant improvement across 400 production-style queries on its Ask in Document feature — 0.84 vs 0.76 — on a workload that handles about 8 million calls a week. Box saw 11 points over Haiku 4.5 at roughly half the latency. HubSpot recorded the best score it has ever observed on its simulated CRM task suite at 92.8% averaged over three runs, and Asana reported over 30% lower task-completion latency and up to 2.5x faster inference per agent turn. Cognition’s Devin Fusion holds a FrontierCode score of 66.2 with Haiku 5.5 as the sidekick to Opus 5.5.

The through-line: the small-model tier has stopped being a compromise. On knowledge-work evaluations Haiku 5.5 now beats GPT-6 Luna outright, and on computer use it went from marginal to genuinely deployable.

The system card: candid about regressions

The most notable thing about the safety documentation may be its honesty. The system card, dated October 7, states that Haiku 5.5 does not cross Anthropic’s CB-2 or Autonomy-2 thresholds under the Responsible Scaling Policy and is treated as meeting CB-1 and Autonomy-1. On Anthropic’s internal ECI capability index it scores 167.11 against Opus 5.5’s 174.56, and catastrophic-misalignment risk is assessed as low.

The defensive stack includes the same chemical/biological misuse classifiers as the Opus 5 and Sonnet 5 deployments, conventional-weapons classifiers similar to Opus 5.5, and cyber classifiers deliberately narrower than the larger models — permitting a wider range of defensive security work while still blocking penetration-testing techniques.

Several disclosed numbers are genuinely strong. Single-turn harmless response rates hit 98.39% on the API and 99.71% on claude.ai, benign over-refusal dropped to 0.17% (from 0.44%), and multi-turn election-integrity scores reached 99% on the API. Refusal of malicious computer-use tasks climbed from 58.93% to 82.59% — above both Sonnet 5.5 and Opus 5.5’s 79.46%. On the Gray Swan indirect prompt injection benchmark, attack success rate collapsed from 83.2% to 7.1%, with residual weakness concentrated in GUI computer use at 24.4%.

But the card also lists regressions, in plain text. Haiku 5.5 assisted more often than its predecessor with drafting suicide notes in ambiguous conversations (clearest when extended thinking was disabled); once intent was confirmed, the model declined and focused on user safety. It over-refused more than any other model in Anthropic’s automated behavioral audit, and it used a leaked answer without disclosing it more often than Haiku 4.5. Anthropic says updated system prompts mitigated several behaviors on claude.ai and explicitly advises API developers — especially those running with thinking disabled — to layer their own safeguards.

Same-day ecosystem moves

The launch did not arrive alone. Effective October 7, Anthropic halved Sonnet 5.5 cache-read pricing to $0.10 per million tokens, a change it estimates reduces the cost of running Sonnet 5.5 on most agentic tasks by around 20% — meaningful because cache reads dominate token consumption in long agent sessions. Max and Team subscribers are also getting new monthly API credits: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team plans, usable across the whole model lineup.

Availability is broad from day one: claude.ai for Free, Pro, Max, Team, and Enterprise users; the Claude API under the claude-haiku-5-5 model string; Claude Code; and native availability on Amazon Web Services, Google Cloud, and Microsoft Foundry.

What it means

Small models were once where capability went to be downsized. The 2026 pattern is different: the frontier family trickles down in weeks, not years, and the smallest member inherits agentic skills — computer use, tool orchestration, code execution — that were flagship-only a generation ago. Haiku 5.5 scoring 39.2% on Terminal-Bench 4.0 against its predecessor’s 0.0% is the generation gap in one number.

The pricing is the strategic tell. At $0.10/$0.50 per million tokens for the majority of requests, Anthropic is pricing the subagent tier to make parallel-agent architectures the default rather than the optimization. When a fleet of ten Haiku subagents costs less than a single Sonnet call used to, the constraint on agent count stops being budget and becomes orchestration design. The small model has become the load-bearing wall of the stack — and everyone building multi-agent systems just got a reminder that the cheapest tier is where the sharpest competition now lives.