Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
Daniel Deutsch, George F. Foster, Markus Freitag
Abstract
Kendall’s tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be handled, a gray area that has motivated different variants in the literature. We demonstrate that, in settings like modern MT meta-evaluation, existing variants have weaknesses arising from their handling of ties, and in some situations can even be gamed. We propose instead to meta-evaluate metrics with a version of pairwise accuracy that gives metrics credit for correctly predicting ties, in combination with a tie calibration procedure that automatically introduces ties into metric scores, enabling fair comparison between metrics that do and do not predict ties. We argue and provide experimental evidence that these modifications lead to fairer ranking-based assessments of metric performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f42ec4d-134a-4603-bc9f-8e55a1fdeed2Cited by top-tier papers32
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan et al.NeurIPS 2025 · 284 citations
- VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video GenerationJiazheng Xu, Yu Huang, Jiale Cheng, Yuanming Yang et al.AAAI 2026 · 112 citations
- Cycle Consistency as Reward: Learning Image-Text Alignment Without Human PreferencesHyojin Bahng, Caroline Chan, Frédo Durand, Phillip IsolaICCV 2025 · 25 citations
- VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement LearningXuanyu Zhang, Weiqi Li, Shijie Zhao, Junlin Li et al.AAAI 2026 · 20 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 14 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
Related papers
- Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 PapersBenjamin Marie, Atsushi Fujita, Raphael RubinoACL 2021
- Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Edoardo Barba et al.ACL 2024
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 13 citations
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba et al.EMNLP 2024 · 1 citation
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
