ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
Shaomu Tan, Christof Monz
Abstract
A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but performs poorly at segment level. In this work, we propose ReMedy, a novel MT metric framework that reformulates translation evaluation as a reward modeling task. Instead of regressing on imperfect human ratings directly, ReMedy learns relative translation quality using pairwise preference data, resulting in a more reliable evaluation. In extensive experiments across WMT22-24 shared tasks (39 language pairs, 111 MT systems), ReMedy achieves stateof-the-art performance at both segment-and system-level evaluation. Specifically, ReMedy-9B surpasses larger WMT winners and massive closed LLMs such as MetricX-13B, XCOMET-Ensemble, GEMBA-GPT-4, PaLM-540B, and finetuned PaLM2. Further analyses demonstrate that ReMedy delivers superior capability in detecting translation errors and evaluating low-quality translations. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary TranslationShaomu Tan, Dawei Zhu, Ke Tran, Michael J. Denkowski et al.ACL 2026 · 1 citation
- M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine TranslationHao Wang, Linlong Xu, Heng Liu, Yangyang Liu et al.ACL 2026
- PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine TranslationLorenzo Proietti, Roman Grundkiewicz, Matt PostACL 2026
Builds on10
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan et al.ICML 2024 · 447 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- Beyond Reference: Evaluating High Quality Translations Better than Human ReferencesKeonwoong Noh, Seokjin Oh, Woohwan JungEMNLP 2024
- What do Large Language Models Need for Machine Translation Evaluation?Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia et al.EMNLP 2024 · 4 citations
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 13 citations
- Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation EvaluationYanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang et al.ACL 2026 · 3 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
