The Best AI Translator Scores 39%: Inside the Last Translation Benchmark
244 researchers built a live benchmark of 3,456 adversarial examples that provably break machine translation. The best model — Gemini 3.1 Pro — passes just 39.3%.
Machine translation has a dirty secret: nobody can actually tell how good the models are anymore. Standard benchmarks are saturated, automatic metrics are unreliable and vulnerable to reward hacking, and even expert human evaluation lacks reproducibility. A coalition of 244 researchers led by Vilém Zouhar (ETH Zurich / Charles University) has now published the most aggressive answer yet — the Last Translation Benchmark (LTB), a live, crowdsourced dataset of 3,456 examples that provably break state-of-the-art translation models, each paired with a handcrafted verification rule that turns embarrassing failures into reproducible tests.
The headline number is brutal. In the benchmark’s main “Blind” evaluation, the single best system in the world — Google’s Gemini 3.1 Pro — successfully translates only 39.30% of the examples. Human performance on the same set is 99.89%. That is not a gap; it is a chasm, and it was measured on September 4, 2026, not in some distant past.
What the Last Translation Benchmark actually is
The paper (arXiv:2609.04173, submitted September 3, 2026) opens with a diagnosis the MT community has quietly accepted for years. As models grew stronger, the classic test sets — newswire-style sentences with a single gold reference — approached saturation. Metrics like BLEU and its learned successors can be gamed; the authors explicitly flag them as “vulnerable to reward-hacking” and “unactionable.” Even paying professional annotators doesn’t fully solve it, because human judgment of translation quality is notoriously non-reproducible between sessions and annotators.
The LTB inverts the usual benchmark design. Instead of sampling representative text, it deliberately collects hard-to-translate inputs across four modalities — text, image, audio, and video — submitted by contributors who must demonstrate that leading translation models provably fail on them. Every submission goes through peer review by other speakers of the relevant languages, and each accepted example ships with a verification rule: a concrete, automatically checkable specification of what a correct translation must (or must not) contain. Think of it as unit tests for translation, written by humans, that any future model can be graded against without a human in the loop.
The scale of the crowdsourcing effort is the paper’s second achievement. LTBv1 — the first frozen release, containing all contributions accepted before September 1, 2026 — lists 244 authors, most of whom earned co-authorship by getting 10 or more submissions through review. The project spans 39 source languages from Algerian Arabic to Vietnamese, targeting 14 languages including Belarusian, Bodo, Croatian, and Swiss German — deliberately including tongues that commercial systems chronically underserve.
The leaderboard nobody wants to top
Because every example carries a machine-checkable pass/fail rule, the team could run a live leaderboard where scores are verifier pass rates, judged using Gemini 3.1 Pro as the verification engine. The Blind-mode results (no privileged information given to the translating model) form a damning snapshot of the field:
| System | Institution | Pass rate |
|---|---|---|
| Human | — | 99.89% |
| Gemini 3.1 Pro | 39.30% | |
| GPT-5.6 Sol | OpenAI | 28.10% |
| GPT-5.6 Luna | OpenAI | 15.37% |
| Gemini 3.5 Flash-Lite | 11.86% | |
| Kimi K3 | Moonshot AI | 11.75% |
| DeepSeek-V4-Pro | DeepSeek | 10.76% |
| Qwen3.7-Plus | Alibaba | 10.21% |
| Google Translate | 0.77% |
Two readings stand out. First, the frontier general-purpose models cluster far ahead of dedicated translation infrastructure — Google Translate, the system that arguably defines machine translation for billions of people, passes less than one example in a hundred on this set. Second, even the best frontier model fails on the majority of cases that humans find merely difficult, not impossible. The benchmark also records an “Oracle” mode, where models get access to the verification rules or human translations — a way to measure headroom when the target is made explicit — though at publication time the Blind table is the populated one.
Zouhar’s own summary of the findings cuts against the usual framing: “Machine translation doesn’t break on just figurative language as one would expect. In fact, for the next generation of models, we may need to invest heavily into multilingual (& cultural) reasoning.” The failures are less about poetry and idioms than about models lacking the world knowledge, cultural context, and deliberate reasoning needed to navigate adversarial inputs — an awkward result for an industry that has largely reframed translation as a solved sub-problem of general intelligence.
Why verification rules matter more than the examples
The deeper contribution is methodological. A benchmark of hard examples alone would eventually saturate like its predecessors. What keeps LTB honest is that each example encodes why models fail: the handcrafted rule states the concrete failure case, so when a future model passes, you know precisely which capability closed the gap — and when it fails, the rule tells you what to fix. That converts evaluation from a leaderboard ritual into an engineering diagnostic, the same shift that verifiable-code benchmarks brought to programming agents.
It also sidesteps the metric crisis. There is no learned metric to hack between the model and its score; there is only a deterministic rule plus a strong verifier model (Gemini 3.1 Pro), and the verifier itself can be swapped as stronger models arrive. Purists will note the circularity risk — using a frontier LLM to judge frontier LLMs — but with humans at 99.89% as an anchor, the verifier’s own blind spots are bounded and visible.
A live paper, not a one-shot
Unlike a static benchmark, the LTB is explicitly a live dataset: contributions are accepted on a rolling basis, and the team plans updated versions of both dataset and paper “with new authors on a rolling basis until end of 2026.” Anyone who gets 10 approved submissions joins the author list — a structure that has already produced one of the largest author collectives in computational linguistics. Submissions are reviewed through a magic-link platform, with reviewers self-nominating by language pair, and the entire pipeline (dataset LTBv1, 90 MB; GitHub repo; leaderboard submissions) is public.
For model developers, the practical takeaway is uncomfortable: your flagship’s 39% is now a public, incrementally updated number, and the only way to move it is to actually fix multilingual and cultural reasoning rather than tune against a static test set. For everyone else, the benchmark is a rare, honest map of where machine translation genuinely stands — and it says the machines are not nearly done.