Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Cassidy Laidlaw, Shivam Singhal, Anca D. Dragan
摘要
Because it is difficult to precisely specify complex objectives, reinforcement learning policies are often optimized using proxy reward functions that only approximate the true goal. However, optimizing proxy rewards frequently leads to reward hacking: the optimized reward function ceases to be a good proxy and the resulting policy performs poorly with respect to the unspecified true reward. Principled solutions to reward hacking have been impeded by the lack of a good definition for the problem. To address this gap, we introduce a definition of reward hacking based on the correlation between proxy and true rewards for states and actions seen by a "reference policy" that breaks down under optimization. We show that this definition captures reward hacking behavior across several realistic settings, including in reinforcement learning from human feedback (RLHF). Using our formulation, we show theoretically that regularization to the reference policy can effectively prevent reward hacking. While the current practice in RLHF applies a KL penalty between action distributions for this purpose, our theory suggests regularizing the χ 2 divergence between the policies' occupancy measures can be more effective. We intuitively show the benefits of this type of regularization and demonstrate that it better mitigates reward hacking in practice across four realistic settings, including RLHF. Our code is available at https://github.com/cassidylaidlaw/orpo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n SamplingLin Gui, Cristina Garbacea, Victor VeitchNeurIPS 2024 · 被引用 138 次
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling 等ICLR 2026 · 被引用 112 次
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju 等NeurIPS 2025 · 被引用 38 次
- Is it Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning EffortXinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell 等ICLR 2026 · 被引用 27 次
- Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable RewardsHieu Trung Nguyen, Bao Nguyen, Wenao Ma, Yuzhi Zhao 等ICLR 2026 · 被引用 19 次
它引用的顶会 Paper34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
相关 Paper
- Robust Optimization for Mitigating Reward Hacking with Correlated ProxiesZixuan Liu, Xiaolin Sun, Zizhan ZhengICLR 2026 · 被引用 2 次
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable RewardsJohannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi SugiyamaICML 2026
- Unifying Stable Optimization and Reference Regularization in RLHFLi He, Qiang Qu, He Zhao, Stephen Wan 等ICLR 2026 · 被引用 6 次
- The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward HackingYuchun Miao, Sen Zhang, Liang Ding, Yuqi Zhang 等ICML 2025
- Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecificationThomas Kwa, Drake Thomas, Adrià Garriga-AlonsoNeurIPS 2024 · 被引用 22 次
