ODIN: Disentangled Reward Mitigates Hacking in RLHF
Lichang Chen, Chen Zhu, Jiuhai Chen, Davit Soselia, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, Bryan Catanzaro
摘要
In this work, we study the issue of reward hacking on the response length, a challenge emerging in Reinforcement Learning from Human Feedback (RLHF) on LLMs. A well-formatted, verbose but less helpful response from the LLMs can often deceive LLMs or even human evaluators to achieve high scores. The same issue also holds for some reward models in RL. To address the challenges in both training and evaluation, we establish a more reliable evaluation protocol for comparing different training configurations, which inspects the trade-off between LLM evaluation score and response length obtained by varying training hyperparameters. Based on this evaluation, we conduct large-scale studies, where the results shed insights into the efficacy of hyperparameters and tricks used in RL on mitigating length bias. We further propose to improve the reward model by jointly training two linear heads on shared feature representations to predict the rewards, one trained to correlate with length, and the other trained to decorrelate with length and therefore focus more on the actual content. We then discard the length head in RL to prevent reward hacking on length. Experiments demonstrate that our approach almost eliminates the reward correlation with length, and improves the obtained policy by a significant margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Self-Consuming Generative Models with Curated Data Provably Optimize Human PreferencesDamien Ferbach, Quentin Bertrand, Avishek Joey Bose, Gauthier GidelNeurIPS 2024 · 被引用 41 次
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju 等NeurIPS 2025 · 被引用 38 次
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference ModelsAnirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi, Mark YatskarICLR 2026 · 被引用 15 次
- Adaptive Batch-Wise Sample Scheduling for Direct Preference OptimizationZixuan Huang, Yikun Ban, Lean Fu, Xiaojie Li 等NeurIPS 2025 · 被引用 14 次
- Bradley-Terry and Multi-Objective Reward Modeling Are ComplementaryZhiwei Zhang, Hui Liu, Xiaomin Li, Zhenwei Dai 等ICLR 2026 · 被引用 8 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
相关 Paper
- Factored Causal Representation Learning for Robust Reward Modeling in RLHFYupei Yang, Lin Yang, Wanxi Deng, Lin Qu 等ICML 2026 · 被引用 1 次
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHFKangwen Zhao, Jianfeng Cai, Jinhua Zhu, Ruopei Sun 等ACL 2026 · 被引用 7 次
- Mitigating Length Bias in RLHF Through a Causal LensHyeonji Kim, Sujeong Oh, Sanghack LeeAAAI 2026 · 被引用 3 次
- Reward Sharpness-Aware Fine-Tuning for Diffusion ModelsKwanyoung Kim, Byeongsu SimCVPR 2026 · 被引用 1 次
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable RewardsJohannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi SugiyamaICML 2026
