Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
- Last Translation Benchmark ships peer-reviewed hard examples, each paired with handcrafted verification rules that flag concrete failure cases instead of assigning a metric score.
- LTBv1 covers submissions accepted before September 1st 2026, spans texts, images, audio and videos, and is designed to accept ongoing contributions.
- The authors argue MT is stuck: standard benchmarks are saturating, automatic metrics are gameable, and human evaluation lacks reproducibility and scalability.