Disentangling Uncertainty in Machine Translation Evaluation
Chrysoula Zerva, Taisiya Glushkova, Ricardo Rei, André F. T. Martins
摘要
Trainable evaluation metrics for machine translation (MT) exhibit strong correlation with human judgements, but they are often hard to interpret and might produce unreliable scores under noisy or out-of-domain data. Recent work has attempted to mitigate this with simple uncertainty quantification techniques (Monte Carlo dropout and deep ensembles), however these techniques (as we show) are limited in several ways -for example, they are unable to distinguish between different kinds of uncertainty, and they are time and memory consuming. In this paper, we propose more powerful and efficient uncertainty predictors for MT evaluation, and we assess their ability to target different sources of aleatoric and epistemic uncertainty. To this end, we develop and compare training objectives for the COMET metric to enhance it with an uncertainty prediction output, including heteroscedastic regression, divergence minimization, and direct uncertainty prediction. Our experiments show improved results on uncertainty prediction for the WMT metrics task datasets, with a substantial reduction in computational costs. Moreover, they demonstrate the ability of these predictors to address specific uncertainty causes in MT evaluation, such as low quality references and outof-domain data. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual RelationshipsAngie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind SatyanarayanCHI 2025 · 被引用 4 次
- Toward Machine Interpreting: Lessons from Human Interpreting StudiesMatthias Sperber, Maureen de Seyssel, Jiajun Bao, Matthias PaulikEMNLP 2025 · 被引用 1 次
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 被引用 273 次
- On the Practicality of Deterministic Epistemic UncertaintyJanis Postels, Mattia Segù, Tao Sun, Luca Daniel Sieber 等ICML 2022 · 被引用 76 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 被引用 13 次
相关 Paper
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 被引用 13 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Test-time Adaptation for Machine Translation Evaluation by Uncertainty MinimizationRunzhe Zhan, Xuebo Liu, Derek F. Wong, Cuilian Zhang 等ACL 2023 · 被引用 1 次
- xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation MetricsDaniil Larionov, Mikhail Seleznyov, Vasiliy Viskov, Alexander Panchenko 等EMNLP 2024 · 被引用 2 次
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song 等EMNLP 2022 · 被引用 18 次
