ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
Shaomu Tan, Christof Monz
摘要
A key challenge in MT evaluation is the inherent noise and inconsistency of human ratings. Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but performs poorly at segment level. In this work, we propose ReMedy, a novel MT metric framework that reformulates translation evaluation as a reward modeling task. Instead of regressing on imperfect human ratings directly, ReMedy learns relative translation quality using pairwise preference data, resulting in a more reliable evaluation. In extensive experiments across WMT22-24 shared tasks (39 language pairs, 111 MT systems), ReMedy achieves stateof-the-art performance at both segment-and system-level evaluation. Specifically, ReMedy-9B surpasses larger WMT winners and massive closed LLMs such as MetricX-13B, XCOMET-Ensemble, GEMBA-GPT-4, PaLM-540B, and finetuned PaLM2. Further analyses demonstrate that ReMedy delivers superior capability in detecting translation errors and evaluating low-quality translations. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary TranslationShaomu Tan, Dawei Zhu, Ke Tran, Michael J. Denkowski 等ACL 2026 · 被引用 1 次
- M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine TranslationHao Wang, Linlong Xu, Heng Liu, Yangyang Liu 等ACL 2026
- PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine TranslationLorenzo Proietti, Roman Grundkiewicz, Matt PostACL 2026
它引用的顶会 Paper10
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan 等ICML 2024 · 被引用 447 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
相关 Paper
- Beyond Reference: Evaluating High Quality Translations Better than Human ReferencesKeonwoong Noh, Seokjin Oh, Woohwan JungEMNLP 2024
- What do Large Language Models Need for Machine Translation Evaluation?Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia 等EMNLP 2024 · 被引用 4 次
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 被引用 13 次
- Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation EvaluationYanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang 等ACL 2026 · 被引用 3 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
