Language Imbalance Driven Rewarding for Multilingual Self-improving
Wen Yang, Junhong Wu, Chen Wang, Chengqing Zong, Jiajun Zhang
Abstract
Large Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such as English and Chinese, leaving many other languages underrepresented. This imbalance, while limiting broader applications, generates a natural preference ranking between languages, offering an opportunity to bootstrap the multilingual capabilities of LLM in a self-improving manner. Thus, we propose Language Imbalance Driven Rewarding, where the inherent imbalance between dominant and non-dominant languages within LLMs is leveraged as a reward signal. Iterative DPO training demonstrates that this approach not only enhances LLM performance in non-dominant languages but also improves the dominant language's capacity, thereby yielding an iterative reward signal. Fine-tuning Meta-Llama-3-8B-Instruct over two iterations of this approach results in continuous improvements in multilingual performance across instructionfollowing and arithmetic reasoning tasks, evidenced by an average improvement of 7.46% win rate on the X-AlpacaEval leaderboard and 13.9% accuracy on the MGSM benchmark. This work serves as an initial exploration, paving the way for multilingual self-improvement of LLMs. The code is available at https: //github.com/ZNLP/Language-Imbalance-Driven-Rewarding .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8908ad3e-385f-472c-9831-9e7a5bcafeacCited by top-tier papers11
- Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement LearningXue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang et al.ACL 2026 · 13 citations
- MPO: Multilingual Safety Alignment via Reward Gap OptimizationWeixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu et al.ACL 2025 · 12 citations
- Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety AlignmentYuyan Bu, Xiaohao Liu, ZhaoXing Ren, Yaodong Yang et al.ICLR 2026 · 9 citations
- Evaluating and Improving Cultural Awareness of Reward Models for LLM AlignmentHongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang et al.ICLR 2026 · 4 citations
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language ModelsPu Jian, Junhong Wu, Wei Sun, Chen Wang et al.EMNLP 2025 · 2 citations
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMsJohn Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer et al.EMNLP 2024 · 4 citations
- Mutual-Taught for Co-adapting Policy and Reward ModelsTianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong et al.ACL 2025 · 1 citation
- MAPO: Advancing Multilingual Reasoning through Multilingual-Alignment-as-Preference OptimizationShuaijie She, Wei Zou, Shujian Huang, Wenhao Zhu et al.ACL 2024
- X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleHaoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang et al.ICLR 2025
