Last Translation Benchmark ships MT hard cases with fail rules
TL;DR
- Last Translation Benchmark ships peer-reviewed hard examples, each paired with handcrafted verification rules that flag concrete failure cases instead of assigning a metric score.
- LTBv1 covers submissions accepted before September 1st 2026, spans texts, images, audio and videos, and is designed to accept ongoing contributions.
- The authors argue MT is stuck: standard benchmarks are saturating, automatic metrics are gameable, and human evaluation lacks reproducibility and scalability.
The Last Translation Benchmark, posted on arxiv, is a collection of examples "that break leading machine translation models," each paired with "handcrafted verification rules describing concrete failure cases." The dataset is designed to stay live and accept ongoing contributions. LTBv1, the current release, covers submissions accepted before September 1st 2026 and spans texts, images, audio and videos.
The paper is blunt about why it exists. The authors write that "standard benchmarks for machine translation are approaching saturation," and that automatic metrics are "unreliable, vulnerable to reward-hacking, and provide unactionable assessments." Even gold human evaluation, the paper says, "often lacks reproducibility, objectivity, and scalability."
Instead of assigning a score, each item defines what a failing translation would look like on that specific example, so a model either trips a rule or it does not.
Led by Vilém Zouhar and Niyati Bafna, the collaboration lists hundreds of contributing authors. Two researchers on our radar had shared the arxiv link within a day of its posting.
Shared on Bluesky by 2 AI experts
-
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: Last Translation Benchmark