MT-Ranker: Reference-free machine translation evaluation by inter-system ranking
Ibraheem Muhammad Moosa, Rui Zhang, Wenpeng Yin
Abstract
Traditionally, Machine Translation (MT) Evaluation has been treated as a regression problem -- producing an absolute translation-quality score. This approach has two limitations: i) the scores lack interpretability, and human annotators struggle with giving consistent scores; ii) most scoring methods are based on (reference, translation) pairs, limiting their applicability in real-world scenarios where references are absent. In practice, we often care about whether a new MT system is better or worse than some competitors. In addition, reference-free MT evaluation is increasingly practical and necessary. Unfortunately, these two practical considerations have yet to be jointly explored. In this work, we formulate the reference-free MT evaluation into a pairwise ranking problem. Given the source sentence and a pair of translations, our system predicts which translation is better. In addition to proposing this new formulation, we further show that this new paradigm can demonstrate superior correlation with human judgments by merely using indirect supervision from natural language inference and weak supervision from our synthetic data. In the context of reference-free evaluation, MT-Ranker, trained without any human annotations, achieves state-of-the-art results on the WMT Shared Metrics Task benchmarks DARR20, MQM20, and MQM21. On a more challenging benchmark, ACES, which contains fine-grained evaluation criteria such as addition, omission, and mistranslation errors, MT-Ranker marks state-of-the-art against reference-free as well as reference-based baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0e8b8b1-1f7f-420e-96e8-e67566050862Cited by top-tier papers5
- xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation MetricsDaniil Larionov, Mikhail Seleznyov, Vasiliy Viskov, Alexander Panchenko et al.EMNLP 2024 · 2 citations
- Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text GenerationTathagata Raha, Clément Christophe, Nada Saadi, Hamza Javed et al.ACL 2026
- ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward ModelingShaomu Tan, Christof MonzEMNLP 2025
- FiRE: Fine-grained Ranking Evaluation for Machine TranslationWenyang Gao, Yinghao Yang, Xi Jin, Jing Li et al.ICML 2026
- PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine TranslationLorenzo Proietti, Roman Grundkiewicz, Matt PostACL 2026
Builds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- BLEU might be Guilty but References are not InnocentMarkus Freitag, David Grangier, Isaac CaswellEMNLP 2020 · 13 citations
- Automatic Machine Translation Evaluation in Many Languages via Zero-Shot ParaphrasingBrian Thompson, Matt PostEMNLP 2020 · 7 citations
Related papers
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba et al.EMNLP 2024 · 1 citation
- On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation EvaluationWei Zhao, Goran Glavas, Maxime Peyrard, Yang Gao et al.ACL 2020 · 53 citations
- Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsNitika Mathur, Timothy Baldwin, Trevor CohnACL 2020 · 14 citations
- Enhancing Human Evaluation in Machine Translation with Comparative JudgementYixiao Song, Parker Riley, Daniel Deutsch, Markus FreitagACL 2025 · 2 citations
- Multi-Hypothesis Machine Translation EvaluationMarina Fomicheva, Lucia Specia, Francisco GuzmánACL 2020 · 13 citations
