Two Models, One Frontier: Sakana's Fugu Max and Fugu Ultra v2 Bet That Orchestration Beats Monoliths
Sakana AI ships Fugu Max (frontier-grade scores at 40-60% lower output cost) and Fugu Ultra v2 (best on 5 of 8 benchmarks — without GPT-6-Astra or Fable 5 in the pool).
For a decade, the AI industry has raced along a single axis: make the foundation model bigger. On September 11, 2026, Tokyo-based Sakana AI pushed hard in a different direction, releasing two orchestrators — Fugu Max and Fugu Ultra v2 — that argue the real frontier is two-dimensional: capability on one axis, cost on the other. The releases are notable not just for the benchmark numbers, but for a provocative design choice in the flagship: Fugu Ultra v2 hits frontier-level scores without any closed frontier model in its agent pool.
The thesis: intelligence is a routing problem
Sakana’s core claim, developed across a rapid five-month cadence of releases, is that “the most capable AI will never come from a single monolithic model.” Instead, capability emerges from a coordination layer that decomposes tasks, routes each piece to the leanest capable model in a swappable pool, and verifies the results. A system that deploys a multi-trillion-parameter model to answer a simple data lookup, the company argues, isn’t intelligent — it’s wasteful.
Fugu Max and Fugu Ultra v2 are the same orchestration architecture tuned for two different missions:
- Fugu Max asks: what is the best possible output at the lowest possible cost?
- Fugu Ultra v2 asks: what is the absolute highest capability on complex, multi-step tasks?
Fugu Max: the Pareto play
Fugu Max orchestrates Sakana’s largest pool of open-weights and specialized models to date — including the NVIDIA Nemotron family, integrated through the partnership Sakana announced in August. The results the company publishes are pointed:
- Best overall score on six benchmarks, including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish (Sakana’s internal benchmark built from its own real-world coding challenges).
- Pricing of $2 per million input tokens and $6 per million output tokens — output pricing the company says is 40–60% lower than Sonnet 5, GPT 5.6 Terra, and Kimi K3.
- Expands the cost-performance Pareto frontier on seven of ten benchmarks, meaning no single model in the comparison set simultaneously matches its quality and price.
In other words, Fugu Max isn’t claiming to be the smartest system available — it’s claiming to occupy a point on the capability-cost curve that no single-model provider can reach, because single models can’t be cheap for easy tasks and brilliant for hard ones at the same time. An orchestrator can, by simply routing easy work to small models.
Fugu Ultra v2: frontier scores, no frontier models
The more striking release is Fugu Ultra v2, tuned for sustained multi-step reasoning, autonomous research, and full-stack software development. The headline numbers:
- Best or joint-best on five of eight benchmarks: GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon, with top-2 placements on seven of eight.
- Chartography (visual reasoning and data interpretation): 48.3, versus Opus 5 at 27.3 and Fable 5 at 29.5 — a gap of nearly 20 points.
- DeepSWE (real-world software engineering): 74.3, outperforming models that cost three to five times more per token.
- Tops SWEFish, demonstrating strength on real-world coding workloads.
The caveat that makes this interesting is buried in the fine print of the benchmark chart: Fugu Ultra v2’s agent pool does not include Fable 5, Fable 5.1, or GPT-6-Astra (its training cutoff is August 28, 2026, and the closed frontier models are excluded from the pool). The scores are achieved by coordinating a swappable pool of open and specialized models. That framing is aimed directly at the two biggest anxieties of enterprise AI buyers in 2026: vendor lock-in and abrupt service cutoffs — both of which the industry has seen play out this year. Sakana pitches this as “AI sovereignty”: supply-chain resilience by design, since any model in the pool can be swapped without re-architecting the system.
From beta thesis to enterprise engine in five months
The release pace itself tells a story about how quickly orchestration has matured as a product category:
- April — beta: proved multi-agent orchestration works as a unified foundation model.
- June — general availability plus Fugu Ultra v1, showing an orchestration layer could match closed frontier models on hard benchmarks.
- July — Fugu-Cyber and a Claude Code interface, specializing orchestration for cybersecurity and coding environments.
- August — Sakana Chat brought orchestration to daily consumer use, and the NVIDIA partnership began integrating Nemotron open models.
- September — Fugu Max and Fugu Ultra v2.
Analysis: the economics may matter more than the benchmarks
Benchmark claims from labs about their own systems deserve the usual skepticism — SWEFish is Sakana’s internal benchmark, and orchestrators benefit structurally from benchmarks that reward decomposition and verification. But two things give this release weight beyond the scores.
First, the economics are auditable in a way benchmarks aren’t. If Fugu Max genuinely delivers output within striking distance of elite models at $6 per million output tokens, the 40–60% cost gap against Sonnet 5, GPT 5.6 Terra, and Kimi K3 compounds brutally at enterprise scale. Routing is a commodity discipline — the hard part is the coordination policy, and Sakana has now iterated it in production across consumer chat, cybersecurity, and coding for five straight months.
Second, the timing is sharp. The release lands in a week when OpenAI paused new ChatGPT Pro sign-ups because demand outstripped capacity, and when regulators and enterprises alike are nervous about dependence on a handful of closed providers. A system whose frontier-grade results provably don’t require the most scarce closed models is an argument that the scarcity itself — and the pricing power it confers — is optional.
The counterargument is equally clear: orchestration trades latency for quality (Fugu Ultra’s own technical report says as much), adds architectural complexity, and multiplies failure modes. For latency-sensitive applications, a single strong model remains simpler. And Sakana’s claim that orchestration “consistently outperforms isolated models” will need to survive independent evaluation, not just the company’s own charts.
Both models are available now via Sakana’s OpenAI-compatible API; existing Fugu users upgrade with a single-line parameter change. The technical report is on arXiv. Whether or not orchestration ultimately displaces the monolithic scaling race, the Pareto frontier just got more crowded — and cheaper.