Enhancing Human Evaluation in Machine Translation with Comparative Judgement
Yixiao Song, Parker Riley, Daniel Deutsch, Markus Freitag
Abstract
Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups-point-wise Multidimensional Quality Metrics (MQM), side-by-side (S×S) MQM, and its simplified version S×S relative ranking (RR). In MQM, annotators mark error spans with categories and severity levels. S×S MQM extends MQM to pairwise error annotation for two translations of the same input, while S×S RR focuses on selecting the better output without labeling errors. Key findings are: (1) the S×S settings achieve higher inter-annotator agreement than MQM; (2) S×S MQM enhances inter-translation error marking consistency compared to MQM by, on average, 38.5% for explicitly compared MT systems and 19.5% for others; (3) all annotation settings return stable system rankings, with S×S RR offering a more efficient alternative to (S×S) MQM; (4) the S×S settings highlight subtle errors overlooked in MQM without altering absolute system evaluations. To spur further research, we release the triply annotated datasets comprising 377 ZhEn and 104 EnDe annotation examples. 1 * Work done during an internship at Google Translate. 1 Data will be available at https://github.com/ google/wmt-mqm-human-evaluation/tree/main/ generalMT2023 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8037edfc-562c-418b-9a10-e39610dbf76eCited by top-tier papers3
- Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-JudgeHamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska et al.ICML 2026 · 2 citations
- FiRE: Fine-grained Ranking Evaluation for Machine TranslationWenyang Gao, Yinghao Yang, Xi Jin, Jing Li et al.ICML 2026
- PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine TranslationLorenzo Proietti, Roman Grundkiewicz, Matt PostACL 2026
Builds on4
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 25 citations
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationDaniel Deutsch, George F. Foster, Markus FreitagEMNLP 2023 · 14 citations
- kNN-LM Does Not Improve Open-ended Text GenerationShufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella et al.EMNLP 2023 · 3 citations
- The Perils of Using Mechanical Turk to Evaluate Open-Ended Text GenerationMarzena Karpinska, Nader Akoury, Mohit IyyerEMNLP 2021 · 3 citations
Related papers
- MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine TranslationParker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni et al.ACL 2026
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba et al.EMNLP 2024 · 1 citation
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 13 citations
- Quantifying the Impact of Translation Errors on Multilingual LLM EvaluationKlaudia Thellmann, Bernhard Stadler, Michael Färber, Jens LehmannACL 2026
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
