FiRE: Fine-grained Ranking Evaluation for Machine Translation
Wenyang Gao, Yinghao Yang, Xi Jin, Jing Li, Yue Zhang
Abstract
Developing reliable machine translation (MT) systems hinges on our ability to distinguish superior translations from inferior ones. However, existing evaluation paradigms, whether limited to coarse overall rankings or misaligned with human preferences, fail to deliver interpretable, fine‑grained feedback in reference‑free settings. We present a Fine-Grained Ranking Evaluation method (FiRE) that leverages off‑the‑shelf large language models to perform criterion‑driven pairwise comparison across three complementary dimensions: faithfulness, fluency, and consistency of style, instead of producing a single holistic judgment. To enable rigorous meta‑evaluation of evaluation paradigms in the absence of any suitable testbed, we construct the first human‑annotated, reference‑free benchmark for fine-grained ranking evaluation, achieving substantial inter‑annotator agreement. Through meta‑evaluation on this benchmark and existing MQM datasets, FiRE demonstrably outperforms regression‑based and error‑analysis metrics in aligning with human comparative judgments, while providing more informative insights into translation quality. Finally, our examination of LLM evaluator biases (position and self-enhancement) and their handling of tied cases offers guidance for more nuanced MT evaluation. Code and benchmark resources are available at https://github.com/wygao8/FiRE-MT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acbd8fee-f7d7-4e11-9e72-e5ee7d531d45Builds on10
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan et al.ICML 2024 · 447 citations
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan et al.ICLR 2024 · 173 citations
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 13 citations
Related papers
- Enhancing Human Evaluation in Machine Translation with Comparative JudgementYixiao Song, Parker Riley, Daniel Deutsch, Markus FreitagACL 2025 · 2 citations
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman et al.EMNLP 2024 · 3 citations
- What do Large Language Models Need for Machine Translation Evaluation?Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia et al.EMNLP 2024 · 4 citations
- Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation EvaluationYanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang et al.ACL 2026 · 3 citations
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
