huggingface.co web signal

Last Translation Benchmark Ships 3,456 Adversarial MT Examples

TL;DR

  • LTBv1 collects 3,456 peer-reviewed, human-authored examples spanning text, images, audio and video, all engineered to break state-of-the-art machine translation systems.
  • Each example ships with handcrafted verification rules describing concrete failure cases, replacing scalar automatic metrics the authors call reward-hackable and unactionable.
  • The authors argue MT failures extend well past figurative language and that next-generation models will need heavier investment in multilingual and cultural reasoning.

The Last Translation Benchmark, released on Hugging Face, ships 3,456 human-authored and peer-reviewed examples engineered to break leading machine translation systems, each paired with handcrafted verification rules describing the concrete failure. The examples span text, images, audio and video, and the project paper is led by Vilém Zouhar with what the authors call a 'massive crowdsourcing effort'.

The pitch is that standard MT benchmarks are saturating and existing evaluation is compromised. Automatic metrics are 'unreliable, vulnerable to reward-hacking, and provide unactionable assessments,' the abstract says, while gold human evaluation 'often lacks reproducibility, objectivity, and scalability.' The verification rules attached to each example are the authors' answer: a per-example failure spec instead of a scalar score.

The finding pushes past the usual idiom-and-metaphor framing. 'Machine translation doesn't break on just figurative language as one would expect,' the paper states, adding that 'for the next generation of models, we may need to invest heavily into multilingual (& cultural) reasoning.'

LTBv1 contains accepted contributions prior to September 1st 2026, and the dataset is described as live, with future releases planned as new data is collected.