83.6 on WMT26 With 25B Active Parameters: Cohere's North Small Translate Outscores DeepL, Google — and Gives the Weights Away
Cohere's first dedicated translation model — a 218B-parameter MoE with only 25B active — posts 83.60 on WMT26 All Languages, beating DeepL NextGen (81.37), Qwen 3.5 397B (81.56) and Google Translate (68.20), runs on two H100s, and ships as open weights.
For years, the machine translation leaderboard had a settled hierarchy: DeepL and Google Translate at the top for dedicated MT, general-purpose LLMs nipping at their heels, and open-weight models a visible tier behind. On September 10, 2026, Cohere took a sledgehammer to that hierarchy. North Small Translate, the company’s first dedicated translation model in its North model family, scored 83.60 on the WMT26 All Languages benchmark — ahead of DeepL NextGen (81.37), Qwen 3.5 397B A17B (81.56), GLM 5.2 FP8 (76.50), Gemma 4 31B (79.46), and Google Translate, which trailed at 68.20.
And then Cohere did the part nobody expected: it published the weights on Hugging Face under a CC BY-NC 4.0 license for research and non-commercial use.
Right-sized, not downsized
The architecture is a study in efficiency. North Small Translate is a mixture-of-experts (MoE) model with 218 billion total parameters but only 25 billion active per token — the same sparse design philosophy Cohere deployed in Command A+ earlier this year. The practical consequences are concrete:
- 16k input / 16k output context — enough to swallow substantial documents in a single call
- 50+ languages supported — 32 high-resource plus 18 additional languages
- Minimum hardware: a single B200 or a pair of H100s at W4A4 quantization, with several near-lossless quantization variants published alongside the full weights
That last bullet is the quiet headline. A model that outperforms proprietary systems from Google and DeepL now runs on hardware that fits under a desk — which is precisely the point.
Built for sovereign AI
Cohere has spent the past two years positioning itself as the sovereign-AI alternative for governments and enterprises uncomfortable shipping their data through American or Chinese hyperscaler APIs. North Small Translate is arguably the sharpest expression of that strategy yet. Translation is one of the most sensitive data categories an organization processes — legal discovery, patents, diplomatic cables, clinical records — and it’s exactly the workload institutions want to keep on-premises. A best-in-class translation model that runs on two H100s inside your own data center removes the last excuse to pipe that text through someone else’s cloud.
The model builds on Cohere’s multilingual lineage: the Tiny Aya family, the widely cited Aya research program, and last year’s Command A Translate. North Small Translate is the first time that accumulated translation expertise has been packaged as a dedicated, purpose-built MT model rather than a capability of a general-purpose LLM.
The regional story is where it gets interesting
Aggregate scores hide the geography. At the regional level, North Small Translate beats Gemma 4 31B outright across Europe — 82.2 vs 73.9, and 82.74 for EU languages in its Agentic configuration — while running essentially even in South Asia (86.2 vs 86.7). More striking: it outperforms DeepL NextGen in every non-European region tested — MENA, South Asia, Southeast Asia, and East Asia — with the widest margins in South Asia and MENA (roughly 8–10 points ahead), a moderate 4–5 point edge in Southeast Asia, and a narrower 1–3 point lead in East Asia, where DeepL’s 85.41 remains formidable.
That pattern reads like a deliberate bet: DeepL’s historic strength is European language pairs, and Cohere went where DeepL isn’t. For any organization whose translation volume skews Asian or Middle Eastern, the open-weight model isn’t just cheaper — it’s measurably better.
The Agentic variant: translation that checks its own work
The 83.60 headline number belongs to the standard model. But Cohere also reports results for North Small Translate (Agentic) — a configuration in which the model hunts for errors in its own draft translations and fixes them before output. That self-correction loop lifts the score to 84.36. On the WMT scoring rubric, where 80–100 means “perfect or with minor errors,” both configurations land in the top band. The gap between the two numbers — nearly a full point from self-review alone — is an early data point for how far agentic scaffolding can push a specialized model without touching its weights.
Throughput and cost: the compounding advantages
Speed: in Cohere’s internal tests on identical hardware, North Small Translate produced 112 output tokens per second at low concurrency versus 81 for Gemma 4 31B, and 39 versus 30 at high concurrency — 30–38% more throughput, or roughly 1.4x.
Long documents, where most translation models quietly collapse: North Small Translate scores 48.9 on Cohere’s long-context evaluation — translating two full book chapters in a single call — more than double Google Translate (21.3) and Gemma 4 31B (19.4), and ahead of every general-purpose LLM tested.
Cost: in the commercial configuration measured by Cohere, the model delivers an 80.1 quality score at $0.000676 per task, averaging 661 tokens. The comparison Cohere chose to highlight is blunt: Gemini 3.1 Pro Preview (high) costs $0.038928 per task — 5,762% more. Qwen 3.5 397B A17B ($0.004525) and Cohere’s own Command A+ ($0.005158) sit in between, at roughly 7–8x the cost.
The RWS partnership
North Small Translate was developed with RWS, the language-services giant whose Language Weaver platform is a fixture in enterprise localization. RWS’s research teams and language experts shaped the model’s real-world translation behavior throughout development — and for enterprises that need commercial licensing, security review, and a full localization platform rather than raw weights, the model is available through Language Weaver. RWS claims relationships with more than 80% of the world’s top 100 brands, which gives Cohere a distribution channel into exactly the customers that spend real money on translation.
Why this matters
Three things make this release more than a benchmark flex. First, the open weights move: the best-performing dedicated MT model in this evaluation is now downloadable, quantizable, and self-hostable, which resets expectations for what governments and enterprises will accept from proprietary MT vendors. Second, the efficiency story — a sub-1T MoE that beats 397B rivals while running on two GPUs — continues the pattern we’ve seen all year of sparse architectures eating dense ones. Third, the sovereign-AI framing: with the EU, Canada, and a growing list of states demanding local control of sensitive AI workloads, a translation model you can air-gap is a product with a waiting market.
The benchmarks were judged by GPT-5.6-Sol, so purists can quibble about LLM-as-judge calibration. But a two-point lead over DeepL, a fifteen-point lead over Google Translate, and a Hugging Face repo anyone can evaluate themselves makes this hard to dismiss. The translation moat, such as it was, just got a lot shallower.