Calibrating Translation Decoding with Quality Estimation on LLMs
Di Wu, Yibin Lei, Christof Monz
Abstract
Neural machine translation (NMT) systems typically employ maximum a posteriori (MAP) decoding to select the highest-scoring translation from the distribution. However, recent evidence highlights the inadequacy of MAP decoding, often resulting in low-quality or even pathological hypotheses as the decoding objective is only weakly aligned with real-world translation quality. This paper proposes to calibrate hypothesis likelihood with translation quality from a distributional view by directly optimizing their Pearson correlation, thereby enhancing decoding effectiveness. With our method, translation with large language models (LLMs) improves substantially after limited training (2K instances per direction). This improvement is orthogonal to those achieved through supervised fine-tuning, leading to substantial gains across a broad range of metrics and human evaluations. This holds even when applied to top-performing translation-specialized LLMs fine-tuned on highquality translation data, such as Tower, or when compared to recent preference optimization methods, like CPO. Moreover, the calibrated translation likelihood can directly serve as a strong proxy for translation quality, closely approximating or even surpassing some state-of-the-art translation quality estimation models, like CometKiwi. Lastly, our in-depth analysis demonstrates that calibration enhances the effectiveness of MAP decoding, thereby enabling greater efficiency in realworld deployment. The resulting state-of-the-art translation model, which covers 10 languages, along with the accompanying code and human evaluation data, has been released: https://github.com/moore3930/calibrating-llm-mt . Translate the following text from English into Chinese. English: Amsterdam is famous for its hundreds of beautiful canals, millions of bicycles, stunning summers, and disgusting winters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc24f5af-a6bf-46e4-8d1b-47fa0c89028eCited by top-tier papers1
Ask how each one uses itBuilds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
Related papers
- Calibrating Sequence likelihood Improves Conditional Language GenerationYao Zhao, Misha Khalman, Rishabh Joshi, Shashi Narayan et al.ICLR 2023 · 38 citations
- MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding MethodsMara Finkelstein, Markus FreitagICLR 2024 · 39 citations
- LLM-SLM Collaborative Framework of Idiomatic Expression GenerationHui Gao, Changhao Song, Peng Zhang, Jing Zhang et al.ACL 2026
- Multilingual Prompting for Improving LLM Generation DiversityQihan Wang, Shidong Pan, Tal Linzen, Emily BlackEMNLP 2025
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 66 citations
