Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
Tom Kocmi, Vilém Zouhar, Christian Federmann, Matt Post
摘要
Ten years ago, a single metric, BLEU, governed progress in machine translation research. For better or worse, there is no such consensus today, and consequently it is difficult for researchers to develop and retain intuitions about metric deltas that drove earlier research and deployment decisions. This paper investigates the "dynamic range" of a number of modern metrics in an effort to provide a collective understanding of the meaning of differences in scores both within and among metrics; in other words, we ask what point difference x in metric y is required between two systems for humans to notice? We conduct our evaluation on a new large dataset, ToShip23, using it to discover deltas at which metrics achieve system-level differences that are meaningful to humans, which we measure by pairwise system accuracy. We additionally show that this method of establishing delta-accuracy is more stable than the standard use of statistical p-values in regards to testset size. Where data size permits, we also explore the effect of metric deltas and accuracy across finer-grained features such as translation direction, domain, and system closeness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan 等ICML 2024 · 被引用 447 次
- QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationGonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei 等NeurIPS 2024 · 被引用 23 次
- Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time MarkersDaniel D'souza, Julia Kreutzer, Adrien Morisot, Ahmet Üstün 等NeurIPS 2025 · 被引用 3 次
- Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware DeferralAntónio Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei 等EMNLP 2025 · 被引用 2 次
- Calibrating Translation Decoding with Quality Estimation on LLMsDi Wu, Yibin Lei, Christof MonzNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper6
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- DEMETR: Diagnosing Evaluation Metrics for TranslationMarzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song 等EMNLP 2022 · 被引用 18 次
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 被引用 14 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
相关 Paper
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
- Sentence-level Aggregation of Lexical Metrics Correlates Stronger with Human Judgements than Corpus-level AggregationPaulo R. Cavalin, Pedro Henrique Domingues, Claudio S. PinhanezAAAI 2025 · 被引用 5 次
- IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesAnanya B. Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan 等ACL 2023 · 被引用 8 次
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 被引用 12 次
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 被引用 13 次
