← All posts / Models

No Frontier Models in the Pool: Sakana's Fugu Max and Fugu Ultra v2 Beat GPT-6 Astra-Class Results by Orchestrating Open Models

Sakana AI's new orchestration models top hard benchmarks like DeepSWE and Chartography using a swappable pool of open and specialized models — no Fable 5, no GPT-6 Astra in the agent pool — while Fugu Max undercuts frontier pricing by 40-60%.

No Frontier Models in the Pool: Sakana's Fugu Max and Fugu Ultra v2 Beat GPT-6 Astra-Class Results by Orchestrating Open Models

For a decade, the AI industry has raced along a single axis: build a bigger, more expensive monolithic model, charge more per token, repeat. Sakana AI, the Tokyo-based lab founded by former Google Brain researchers, has just shipped its most aggressive argument yet that this axis is the wrong one. On September 11, the company released Fugu Max and Fugu Ultra v2 — two orchestration models that coordinate a swappable pool of open-weights and specialized models behind a single OpenAI-compatible API, and that now sit at or beyond the cost-performance frontier that individual frontier models define.

The detail that should stop everyone in the industry: Fugu Ultra v2 achieves its top scores without Claude Fable 5, Fable 5.1, or GPT-6 Astra anywhere in its agent pool. The best results come purely from orchestrating open and specialized models. In a year when enterprises are increasingly worried about vendor lock-in, API revocations, and sudden model deprecations, that is not a benchmark footnote — it is a procurement argument.

Two models, one architecture

Fugu Max and Fugu Ultra v2 are not separate products so much as the same core orchestration engine optimized for two different missions. Fugu Max asks: what is the best possible output we can deliver at the lowest possible cost? Fugu Ultra v2 asks: what is the absolute highest capability we can achieve on complex, multi-step tasks?

The lineage is short but fast. Fugu debuted in beta in April 2026, proving that multi-agent orchestration could work as a unified foundation model. June brought general availability and Fugu Ultra v1, which matched closed frontier models on hard benchmarks. July added Fugu-Cyber and a Claude Code interface, specializing orchestration for cybersecurity and coding workflows. August saw Sakana Chat bring the system to daily consumer use, alongside a partnership with NVIDIA that began integrating the Nemotron family of open models. September’s release is the culmination: Fugu Max expands the pool of orchestratable models to its largest ever, while Fugu Ultra v2 pushes peak performance higher — explicitly without relying on the frontier models it orchestrates around.

The numbers

Fugu Max is the cost-efficiency play. It achieves the best overall score on six benchmarks — Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish (Sakana’s internal benchmark built on its own coding challenges). At $2 per million input tokens and $6 per million output tokens, its output pricing runs 40-60% lower than Sonnet 5, GPT 5.6 Terra, and Kimi K3. Across ten benchmarks, it expands the cost-performance Pareto frontier on seven. The pitch is blunt: performance within striking distance of elite models at two to six times lower cost.

Fugu Ultra v2 is the ceiling-raiser. It posts the best or joint-best score on five of eight hard benchmarks — GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon — and lands in the top two on seven of eight. Two results stand out:

  • Chartography (visual reasoning and data interpretation): Fugu Ultra v2 scores 48.3, against Opus 5 at 27.3 and Fable 5 at 29.5 — a gap large enough to be a category difference, not a rounding error.
  • DeepSWE (real-world software engineering): 74.3, outperforming models that cost three to five times more per token.

One methodological note that deserves credit: Sakana discloses that Fugu Ultra v2’s training cutoff is August 28, 2026, and that Fable 5, Fable 5.1, and GPT-6 Astra are not in the model pool. Benchmark contamination debates have haunted the industry all year; explicitly stating what the orchestrator can and cannot see is the kind of transparency other labs should copy.

Why orchestration, why now

Three currents in 2026 make this release more than an academic curiosity.

First, open models have become the fastest-growing, most diverse part of the ecosystem — as Sakana itself puts it — and they become dramatically more useful when orchestrated together than when used in isolation. The frontier of capability-per-dollar is increasingly being built out of many open, specialized models working in concert, not one giant model doing everything.

Second, supply-chain resilience has become a board-level concern. The summer’s geopolitical turbulence, API revocations, and the ongoing US-China model export fight have made “what happens if our model vendor cuts us off” a real question in enterprise architecture reviews. A swappable pool of agents — where any single model can be replaced without rearchitecting — is a structural answer. Sakana markets this as “AI sovereignty,” and for once the marketing term maps to a genuine technical property.

Third, the economics are starting to bite. With AI infrastructure spending drawing scrutiny from investors and analysts worried about a pacing bubble, a system that routes each task to the leanest capable model — instead of deploying a multi-trillion-parameter model to execute a simple data lookup — addresses the waste that critics keep flagging. As the release notes dryly observe, a system that does the latter “is not intelligent, but wasteful.”

The catch

Honest caveats apply. Orchestration means more moving parts, more latency variability, and dependency on the quality of routing decisions — a misrouted task lands on a model that cannot handle it. Sakana’s benchmarks are largely its own selections, and SWEFish is an internal benchmark; independent verification will take time. The company is also small relative to the hyperscalers it is positioning against, and orchestration layers are easy for the big labs to imitate if the approach proves out — OpenAI, Google, and Anthropic all have router-like features in various stages of shipping.

But that is precisely the point: the idea is winning even if the company is not the only beneficiary. If the frontier labs adopt orchestration internally, Sakana’s thesis is validated; if they do not, Sakana keeps extending the Pareto frontier in public. Either way, the industry’s unit of competition is shifting from “the model” toward “the system that manages the models.”

Availability

Both models are available now via Sakana’s OpenAI-compatible API. Existing Fugu users upgrade with a single-line parameter change — no migration, same API surface. For teams that have spent 2026 watching their model vendors consolidate, deprecate, and reprice, that stability is part of the product.

The pufferfish has historically been a defensive creature — it survives by being more trouble to swallow than it is worth. As a metaphor for a small lab surviving among giants by making itself architecturally indispensable, it is surprisingly apt.