Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
Thomas Kwa, Drake Thomas, Adrià Garriga-Alonso
Abstract
When applying reinforcement learning from human feedback (RLHF), the reward is learned from data and, therefore, always has some error. It is common to mitigate this by regularizing the policy with KL divergence from a base model, with the hope that balancing reward with regularization will achieve desirable outcomes despite this reward misspecification. We show that when the reward function has light-tailed error, optimal policies under less restrictive KL penalties achieve arbitrarily high utility. However, if error is heavy-tailed, some policies obtain arbitrarily high reward despite achieving no more utility than the base model--a phenomenon we call catastrophic Goodhart. We adapt a discrete optimization method to measure the tails of reward models, finding that they are consistent with light-tailed error. However, the pervasiveness of heavy-tailed distributions in many real-world applications indicates that future sources of RL reward could have heavy-tailed error, increasing the likelihood of reward hacking even with KL regularization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju et al.NeurIPS 2025 · 38 citations
- Unifying Stable Optimization and Reference Regularization in RLHFLi He, Qiang Qu, He Zhao, Stephen Wan et al.ICLR 2026 · 6 citations
- MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiencyNicolas Dufour, Lucas Degeorge, Arijit Ghosh, Vicky Kalogeiton et al.ICML 2026 · 2 citations
- The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low RegretLukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré et al.ICML 2025
- Correlated Proxies: A New Definition and Improved Mitigation for Reward HackingCassidy Laidlaw, Shivam Singhal, Anca D. DraganICLR 2025
Builds on2
Related papers
- Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian AlignerKazusato Oko, Annie Ulichney, Nika Haghtalab, Han BaoICML 2026
- On Teacher Hacking in Language Model DistillationDaniil Tiapkin, Daniele Calandriello, Johan Ferret, Sarah Perrin et al.ICML 2025
- Beyond Expectations: Quantile-Guided Alignment for Risk-Calibrated Language ModelsXinran Wang, Jin Du, Azal Ahmad Khan, Qi Le et al.NeurIPS 2025 · 1 citation
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable RewardsJohannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi SugiyamaICML 2026
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang et al.NeurIPS 2024 · 157 citations
