Defining and Characterizing Reward Gaming
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David Krueger
Abstract
We provide the first formal definition of reward gaming, a phenomenon where optimizing an imperfect proxy reward function, R, leads to poor performance according to a true reward function, R. We say that a proxy is ungameable if increasing the expected proxy return can never decrease the expected true return. Intuitively, it should be possible to create an ungameable proxy by overlooking fine-grained distinctions between roughly equivalent outcomes, but we show this is usually not the case. A key insight is that the linearity of reward (as a function of state-action visit counts) makes ungameability a very strong condition. In particular, for the set of all stochastic policies, two reward functions can only be ungameable if one of them is constant. We thus turn our attention to deterministic policies and finite sets of stochastic policies, where non-trivial ungameable pairs always exist, and establish necessary and sufficient conditions for the existence of simplifications, an important special case of ungameability. Our results reveal a tension between using reward functions to specify narrow tasks and aligning AI systems with human values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e82b84b-2869-419b-8f18-1126c53409faCited by top-tier papers150
- Normalized Rewards for Preference OptimizationShawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald et al.ICML 2026 · 571 citations
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang et al.NeurIPS 2025 · 310 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- Reward Model Ensembles Help Mitigate OveroptimizationThomas Coste, Usman Anwar, Robert Kirk, David KruegerICLR 2024 · 208 citations
Builds on4
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- Optimal Policies Tend To Seek PowerAlexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch et al.NeurIPS 2021 · 111 citations
Related papers
- Correlated Proxies: A New Definition and Improved Mitigation for Reward HackingCassidy Laidlaw, Shivam Singhal, Anca D. DraganICLR 2025
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer et al.ICLR 2024 · 22 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Robust Optimization for Mitigating Reward Hacking with Correlated ProxiesZixuan Liu, Xiaolin Sun, Zizhan ZhengICLR 2026 · 2 citations
- When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPsJose Aguilar Escamilla, Haoyang Hong, Jiawei Li, Haoyu Zhao et al.ICML 2026
