Settling the Reward Hypothesis
Michael Bowling, John D. Martin, David Abel, Will Dabney
2023Year
47Citations
16Top-tier citations
Abstract
The reward hypothesis posits that, "all of what we mean by goals and purposes can be well thought of as maximization of the expected value of the cumulative sum of a received scalar signal (reward)." We aim to fully settle this hypothesis. This will not conclude with a simple affirmation or refutation, but rather specify completely the implicit requirements on goals and purposes under which the hypothesis holds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3df32c5-fed3-42ce-adeb-fea1d93e7103Cited by top-tier papers16
- A Definition of Continual Reinforcement LearningDavid Abel, André Barreto, Benjamin Van Roy, Doina Precup et al.NeurIPS 2023 · 167 citations
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu et al.ICLR 2024 · 97 citations
- Rethinking Decision Transformer via Hierarchical Reinforcement LearningYi Ma, Jianye Hao, Hebin Liang, Chenjun XiaoICML 2024 · 15 citations
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 13 citations
- Plasticity as the Mirror of EmpowermentDavid Abel, Michael Bowling, André Barreto, Will Dabney et al.NeurIPS 2025 · 9 citations
Builds on2
Related papers
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 4 citations
- Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPsMehran Shakerinava, Siamak Ravanbakhsh, Adam M. ObermanNeurIPS 2025 · 3 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer et al.ICLR 2024 · 22 citations
