Enhancing Machine Translation with Self-Supervised Preference Data
Haoxiang Sun, Ruize Gao, Pei Zhang, Baosong Yang, Rui Wang
Abstract
Model alignment methods like Direct Preference Optimization (Rafailov et al., 2024) and Contrastive Preference Optimization (Xu et al., 2024b) have enhanced machine translation performance by leveraging preference data to enable models to reject suboptimal outputs. During preference data construction, previous approaches primarily rely on humans, strong models like GPT4 (OpenAI, 2023) or model self-sampling. In this study, we first explain the shortcomings of this practice. Then, we propose Self-Supervised Preference Optimization (SSPO), a novel framework which efficiently constructs translation preference data for iterative DPO training. Applying SSPO to 14B parameters large language models (LLMs) achieves comparable or better performance than GPT-4o on FLO-RES and multi-domain test datasets. We release an augmented MQM dataset in https: //github.com/sunny-sjtu/MQM-aug . * Work done during internship at Tongyi Lab. † Rui Wang and Baosong Yang are co-corresponding authors. * We use gpt-4o-0806 available from the OpenAI API.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97fcf5a4-609d-4865-a277-c6f328355827Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan et al.ICML 2024 · 447 citations
Related papers
- Adaptive Batch-Wise Sample Scheduling for Direct Preference OptimizationZixuan Huang, Yikun Ban, Lean Fu, Xiaojie Li et al.NeurIPS 2025 · 14 citations
- Finding the Sweet Spot: Preference Data Construction for Scaling Preference OptimizationYao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng et al.ACL 2025 · 8 citations
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen et al.ACL 2026 · 2 citations
- Word Alignment as Preference for Machine TranslationQiyu Wu, Masaaki Nagata, Zhongtao Miao, Yoshimasa TsuruokaEMNLP 2024 · 4 citations
- Self-Play Preference Optimization for Language Model AlignmentYue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji et al.ICLR 2025
