Disentangling Uncertainty in Machine Translation Evaluation
Chrysoula Zerva, Taisiya Glushkova, Ricardo Rei, André F. T. Martins
Abstract
Trainable evaluation metrics for machine translation (MT) exhibit strong correlation with human judgements, but they are often hard to interpret and might produce unreliable scores under noisy or out-of-domain data. Recent work has attempted to mitigate this with simple uncertainty quantification techniques (Monte Carlo dropout and deep ensembles), however these techniques (as we show) are limited in several ways -for example, they are unable to distinguish between different kinds of uncertainty, and they are time and memory consuming. In this paper, we propose more powerful and efficient uncertainty predictors for MT evaluation, and we assess their ability to target different sources of aleatoric and epistemic uncertainty. To this end, we develop and compare training objectives for the COMET metric to enhance it with an uncertainty prediction output, including heteroscedastic regression, divergence minimization, and direct uncertainty prediction. Our experiments show improved results on uncertainty prediction for the WMT metrics task datasets, with a substantial reduction in computational costs. Moreover, they demonstrate the ability of these predictors to address specific uncertainty causes in MT evaluation, such as low quality references and outof-domain data. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual RelationshipsAngie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind SatyanarayanCHI 2025 · 4 citations
- Toward Machine Interpreting: Lessons from Human Interpreting StudiesMatthias Sperber, Maureen de Seyssel, Jiajun Bao, Matthias PaulikEMNLP 2025 · 1 citation
Builds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- On the Practicality of Deterministic Epistemic UncertaintyJanis Postels, Mattia Segù, Tao Sun, Luca Daniel Sieber et al.ICML 2022 · 76 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 13 citations
Related papers
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 13 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- Test-time Adaptation for Machine Translation Evaluation by Uncertainty MinimizationRunzhe Zhan, Xuebo Liu, Derek F. Wong, Cuilian Zhang et al.ACL 2023 · 1 citation
- xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation MetricsDaniil Larionov, Mikhail Seleznyov, Vasiliy Viskov, Alexander Panchenko et al.EMNLP 2024 · 2 citations
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song et al.EMNLP 2022 · 18 citations
