Utility Theory for Sequential Decision Making
Mehran Shakerinava, Siamak Ravanbakhsh
Abstract
The von Neumann-Morgenstern (VNM) utility theorem shows that under certain axioms of rationality, decision-making is reduced to maximizing the expectation of some utility function. We extend these axioms to increasingly structured sequential decision making settings and identify the structure of the corresponding utility functions. In particular, we show that memoryless preferences lead to a utility in the form of a per transition reward and multiplicative factor on the future return. This result motivates a generalization of Markov Decision Processes (MDPs) with this structure on the agent's returns, which we call Affine-Reward MDPs. A stronger constraint on preferences is needed to recover the commonly used cumulative sum of scalar rewards in MDPs. A yet stronger constraint simplifies the utility function for goal-seeking agents in the form of a difference in some function of states that we call potential functions. Our necessary and sufficient conditions demystify the reward hypothesis that underlies the design of rational agents in reinforcement learning by adding an axiom to the VNM rationality axioms and motivates new directions for AI research involving sequential decision making.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 47 citations
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 13 citations
- How does Inverse RL Scale to Large State Spaces? A Provably Efficient ApproachFilippo Lazzati, Mirco Mutti, Alberto Maria MetelliNeurIPS 2024 · 5 citations
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm et al.ICLR 2024 · 3 citations
- Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPsMehran Shakerinava, Siamak Ravanbakhsh, Adam M. ObermanNeurIPS 2025 · 3 citations
Builds on1
Related papers
- Optimal Policies Tend To Seek PowerAlexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch et al.NeurIPS 2021 · 111 citations
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell et al.ICML 2024 · 44 citations
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 96 citations
- Apparently Irrational Choice as Optimal Sequential Decision MakingHaiyang Chen, Hyung Jin Chang, Andrew HowesAAAI 2021 · 10 citations
- An Analytical Study of Utility Functions in Multi-Objective Reinforcement LearningManel Rodriguez-Soto, Juan A. Rodríguez-Aguilar, Maite López-SánchezNeurIPS 2024 · 9 citations
