Last Translation Benchmark · 2026-09-03 · notable
Last Translation Benchmark — 3,456 examples that break translation models
The Last Translation Benchmark is an open set of 3,456 human-written examples that leading machine translation models get wrong. Each example ships with handcrafted rules stating exactly what a correct translation has to do.

A live, peer-reviewed collection of translation examples that today's best models provably fail.
Key specs
| Version | LTBv1 |
|---|---|
| License | CC-BY-4.0 |
| Examples | 3,456 |
What is it?
The Last Translation Benchmark gathers 3,456 human-authored examples — texts plus images, audio and video — picked because leading machine translation models break on them. Release LTBv1 holds every contribution accepted before 1 September 2026, and the collection stays open: the project keeps taking submissions through the end of 2026.
How does it work?
Rather than scoring output with an automatic metric, each example in the Last Translation Benchmark carries handcrafted verification rules that spell out the concrete failure to look for. The public leaderboard reports a verifier pass rate judged by Gemini 3.1 Pro, and runs two modes: a blind mode where the model gets no privileged information, and an oracle mode that hands it the rules and a human translation.
Why does it matter?
Standard machine translation benchmarks are close to saturated, and the paper argues automatic metrics are unreliable and open to reward-hacking while human evaluation is hard to reproduce at scale. Rule-based verification turns a score into a pass or fail a person can re-check, so a miss points at a named defect instead of a number nobody can act on.
Who is it for?
machine translation researchers and evaluation teams
Try it
https://huggingface.co/datasets/zouhar/last-translation-benchmark