Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers
Benjamin Marie, Atsushi Fujita, Raphael Rubino
Abstract
This paper presents the first large-scale metaevaluation of machine translation (MT). We annotated MT evaluations conducted in 769 research papers published from 2010 to 2020. Our study shows that practices for automatic MT evaluation have dramatically changed during the past decade and follow concerning trends. An increasing number of MT evaluations exclusively rely on differences between BLEU scores to draw conclusions, without performing any kind of statistical significance testing nor human evaluation, while at least 108 metrics claiming to be better than BLEU have been proposed. MT evaluations in recent papers tend to copy and compare automatic metric scores from previous work to claim the superiority of a method or an algorithm without confirming neither exactly the same training, validating, and testing data have been used nor the metric scores are comparable. Furthermore, tools for reporting standardized metric scores are still far from being widely adopted by the MT community. After showing how the accumulation of these pitfalls leads to dubious evaluation, we propose a guideline to encourage better automatic MT evaluation along with a simple meta-evaluation scoring method to assess its credibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95f6a68e-7f98-466e-9833-6185e5c4cbd3Cited by top-tier papers16
- Few-shot Controllable Style Transfer for Low-Resource Multilingual SettingsKalpesh Krishna, Deepak Nathani, Xavier Garcia, Bidisha Samanta et al.ACL 2022 · 28 citations
- Modeling the Machine Learning MultiverseSamuel J. Bell, Onno Kampman, Jesse Dodge, Neil D. LawrenceNeurIPS 2022 · 23 citations
- On Robust Prefix-Tuning for Text ClassificationZonghan Yang, Yang LiuICLR 2022 · 23 citations
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song et al.EMNLP 2022 · 18 citations
- Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic FactorsMarvin Kaster, Wei Zhao, Steffen EgerEMNLP 2021 · 13 citations
Builds on2
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningXiaomian Kang, Yang Zhao, Jiajun Zhang, Chengqing ZongEMNLP 2020 · 61 citations
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 14 citations
Related papers
- Navigating the Metrics Maze: Reconciling Score Magnitudes and AccuraciesTom Kocmi, Vilém Zouhar, Christian Federmann, Matt PostACL 2024 · 5 citations
- Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level AggregationPaulo R. Cavalin, Pedro Henrique Domingues, Claudio S. PinhanezAAAI 2025 · 5 citations
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationDaniel Deutsch, George F. Foster, Markus FreitagEMNLP 2023 · 14 citations
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 13 citations
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 13 citations
