Enhancing Human Evaluation in Machine Translation with Comparative Judgement
Yixiao Song, Parker Riley, Daniel Deutsch, Markus Freitag
摘要
Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups-point-wise Multidimensional Quality Metrics (MQM), side-by-side (S×S) MQM, and its simplified version S×S relative ranking (RR). In MQM, annotators mark error spans with categories and severity levels. S×S MQM extends MQM to pairwise error annotation for two translations of the same input, while S×S RR focuses on selecting the better output without labeling errors. Key findings are: (1) the S×S settings achieve higher inter-annotator agreement than MQM; (2) S×S MQM enhances inter-translation error marking consistency compared to MQM by, on average, 38.5% for explicitly compared MT systems and 19.5% for others; (3) all annotation settings return stable system rankings, with S×S RR offering a more efficient alternative to (S×S) MQM; (4) the S×S settings highlight subtle errors overlooked in MQM without altering absolute system evaluations. To spur further research, we release the triply annotated datasets comprising 377 ZhEn and 104 EnDe annotation examples. 1 * Work done during an internship at Google Translate. 1 Data will be available at https://github.com/ google/wmt-mqm-human-evaluation/tree/main/ generalMT2023 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-JudgeHamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska 等ICML 2026 · 被引用 2 次
- FiRE: Fine-grained Ranking Evaluation for Machine TranslationWenyang Gao, Yinghao Yang, Xi Jin, Jing Li 等ICML 2026
- PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine TranslationLorenzo Proietti, Roman Grundkiewicz, Matt PostACL 2026
它引用的顶会 Paper4
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 被引用 25 次
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationDaniel Deutsch, George F. Foster, Markus FreitagEMNLP 2023 · 被引用 14 次
- kNN-LM Does Not Improve Open-ended Text GenerationShufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella 等EMNLP 2023 · 被引用 3 次
- The Perils of Using Mechanical Turk to Evaluate Open-Ended Text GenerationMarzena Karpinska, Nader Akoury, Mohit IyyerEMNLP 2021 · 被引用 3 次
相关 Paper
- MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine TranslationParker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni 等ACL 2026
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba 等EMNLP 2024 · 被引用 1 次
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 被引用 13 次
- Quantifying the Impact of Translation Errors on Multilingual LLM EvaluationKlaudia Thellmann, Bernhard Stadler, Michael Färber, Jens LehmannACL 2026
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
