WPO: Enhancing RLHF with Weighted Preference Optimization
Wenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu
Abstract
Reinforcement learning from human feedback (RLHF) is a promising solution to align large language models (LLMs) more closely with human values. Off-policy preference optimization, where the preference data is obtained from other models, is widely adopted due to its cost efficiency and scalability. However, off-policy preference optimization often suffers from a distributional gap between the policy used for data collection and the target policy, leading to suboptimal optimization. In this paper, we propose a novel strategy to mitigate this problem by simulating on-policy learning with off-policy preference data. Our Weighted Preference Optimization (WPO) method adapts off-policy data to resemble on-policy data more closely by reweighting preference pairs according to their probability under the current policy. This method not only addresses the distributional gap problem but also enhances the optimization process without incurring additional costs. We validate our method on instruction following benchmarks including Alpaca Eval 2 and MT-bench. WPO not only outperforms Direct Preference Optimization (DPO) by up to 5.6% on Alpaca Eval 2 but also establishes a remarkable length-controlled winning rate against GPT-4-turbo of 76.7% based on Gemma-2-9b-it. We release the code and models at https://github.com/wzhouad/WPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cbb8ef4-500a-4c5d-bd9f-2f26f9dd45bfCited by top-tier papers7
- TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMsRicardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves et al.ACL 2026 · 34 citations
- Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMsShangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang et al.ICLR 2026 · 10 citations
- When Weak LLMs Speak with Confidence, Preference Alignment Gets StrongerAmirabbas Afzali, Myeongho Jeon, Maria BrbicICLR 2026
- D: Dynamic Directional Graph-Constrained Data Scheduling for LLM TrainingYuanjian Xu, Jianing Hao, Guang Zhang, Zhong LiICML 2026
- Advantage-Guided Distillation for Preference Alignment in Small Language ModelsShiping Gao, Fanqi Wan, Jiajian Guo, Xiaojun Quan et al.ICLR 2025
Builds on13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
Related papers
- Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay PerspectiveRuichen Shao, Bei Li, Gangao Liu, Yang Chen et al.ICLR 2025
- Self-Play Preference Optimization for Language Model AlignmentYue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji et al.ICLR 2025
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference OptimizationJunming Yang, Ning Xu, Biao Liu, Shiqi Qiao et al.ICLR 2026 · 3 citations
- DPO Meets PPO: Reinforced Token Optimization for RLHFHan Zhong, Zikang Shan, Guhao Feng, Wei Xiong et al.ICML 2025
- Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference ModelJunshu Pan, Wei Shen, Shulin Huang, Qiji Zhou et al.AAAI 2026 · 7 citations
