arxiv.org web signal

Last Translation Benchmark ships MT hard cases with fail rules

TL;DR

  • Last Translation Benchmark ships peer-reviewed hard examples, each paired with handcrafted verification rules that flag concrete failure cases instead of assigning a metric score.
  • LTBv1 covers submissions accepted before September 1st 2026, spans texts, images, audio and videos, and is designed to accept ongoing contributions.
  • The authors argue MT is stuck: standard benchmarks are saturating, automatic metrics are gameable, and human evaluation lacks reproducibility and scalability.

The Last Translation Benchmark, posted on arxiv, is a collection of examples "that break leading machine translation models," each paired with "handcrafted verification rules describing concrete failure cases." The dataset is designed to stay live and accept ongoing contributions. LTBv1, the current release, covers submissions accepted before September 1st 2026 and spans texts, images, audio and videos.

The paper is blunt about why it exists. The authors write that "standard benchmarks for machine translation are approaching saturation," and that automatic metrics are "unreliable, vulnerable to reward-hacking, and provide unactionable assessments." Even gold human evaluation, the paper says, "often lacks reproducibility, objectivity, and scalability."

Instead of assigning a score, each item defines what a failing translation would look like on that specific example, so a model either trips a rule or it does not.

Led by Vilém Zouhar and Niyati Bafna, the collaboration lists hundreds of contributing authors. Two researchers on our radar had shared the arxiv link within a day of its posting.

Shared on Bluesky by 2 AI experts