Contrastive ESA scores multiple MT outputs side by side
TL;DR
- cESA shows annotators several translations of the same source at once, has them mark major and minor error spans, then score each 0-100%.
- The method was validated on an English-to-Japanese comparison of 12 models, with the authors reporting lower annotation time and noise than pointwise evaluation.
- Unlike contrastive ranking, cESA yields absolute quality judgments that support non-parametric model rankings without post-hoc statistical corrections.
A new machine-translation evaluation protocol asks annotators to look at several candidate translations of the same source together rather than one at a time, and its authors report faster, less noisy scoring than the pointwise baseline. In Contrastive ESA: Human Evaluation of Multiple Translations at Once, Vilém Zouhar and colleagues describe cESA, a protocol in which the annotator marks major and minor error spans across the displayed outputs, then assigns each translation a score "from 0% to 100% on absolute scale."
The method was validated on an English-to-Japanese evaluation of 12 models. The abstract states cESA delivers "reductions in annotation time and noise compared to standard pointwise evaluation," though it does not put a figure on either. The paper's pitch against the existing family of contrastive ranking methods is that cESA "yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections."
The preprint was circulating among a couple of the MT evaluation researchers we follow within a day of posting.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Contrastive ESA: Human Evaluation of Multiple Translations at Once