Unifying Stable Optimization and Reference Regularization in RLHF
Li He, Qiang Qu, He Zhao, Stephen Wan, Dadong Wang, Lina Yao, Tongliang Liu
Abstract
Reinforcement Learning from Human Feedback (RLHF) has advanced alignment capabilities significantly but remains hindered by two core challenges: reward hacking and stable optimization. Current solutions independently address these issues through separate regularization strategies, specifically a KL-divergence penalty against a supervised fine-tuned model (π 0 ) to mitigate reward hacking, and policy ratio clipping towards the current policy (π t ) to promote stable alignment. However, the implicit trade-off arising from simultaneously regularizing towards both π 0 and π t remains under-explored. In this paper, we introduce a unified regularization approach that explicitly balances the objectives of preventing reward hacking and maintaining stable policy updates. Our simple yet principled alignment objective yields a weighted supervised fine-tuning loss with a superior trade-off, which demonstrably improves both alignment results and implementation complexity. Extensive experiments across diverse benchmarks validate that our method consistently outperforms RLHF and online preference learning methods, achieving enhanced alignment performance and stability. Our implementation is available at github.com/tmllab/2026 ICLR DAR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60d85ff1-3b5b-47c0-9057-41d0d5d0d70fCited by top-tier papers1
Ask how each one uses itBuilds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Correlated Proxies: A New Definition and Improved Mitigation for Reward HackingCassidy Laidlaw, Shivam Singhal, Anca D. DraganICLR 2025
- Scaling Laws for Reward Model Overoptimization in Direct Alignment AlgorithmsRafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi et al.NeurIPS 2024 · 169 citations
- Decoding-time Realignment of Language ModelsTianlin Liu, Shangmin Guo, Leonardo Bianco, Daniele Calandriello et al.ICML 2024 · 69 citations
- Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference OptimizationAudrey Huang, Wenhao Zhan, Tengyang Xie, Jason D. Lee et al.ICLR 2025
- Joint Reward and Policy Learning with Demonstrations and Human Feedback Improves AlignmentChenliang Li, Siliang Zeng, Zeyi Liao, Jiaxiang Li et al.ICLR 2025
