Utility Theory for Sequential Decision Making
Mehran Shakerinava, Siamak Ravanbakhsh
摘要
The von Neumann-Morgenstern (VNM) utility theorem shows that under certain axioms of rationality, decision-making is reduced to maximizing the expectation of some utility function. We extend these axioms to increasingly structured sequential decision making settings and identify the structure of the corresponding utility functions. In particular, we show that memoryless preferences lead to a utility in the form of a per transition reward and multiplicative factor on the future return. This result motivates a generalization of Markov Decision Processes (MDPs) with this structure on the agent's returns, which we call Affine-Reward MDPs. A stronger constraint on preferences is needed to recover the commonly used cumulative sum of scalar rewards in MDPs. A yet stronger constraint simplifies the utility function for goal-seeking agents in the form of a difference in some function of states that we call potential functions. Our necessary and sufficient conditions demystify the reward hypothesis that underlies the design of rational agents in reinforcement learning by adding an axiom to the VNM rationality axioms and motivates new directions for AI research involving sequential decision making.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 被引用 47 次
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 被引用 13 次
- How does Inverse RL Scale to Large State Spaces? A Provably Efficient ApproachFilippo Lazzati, Mirco Mutti, Alberto Maria MetelliNeurIPS 2024 · 被引用 5 次
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm 等ICLR 2024 · 被引用 3 次
- Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPsMehran Shakerinava, Siamak Ravanbakhsh, Adam M. ObermanNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Optimal Policies Tend To Seek PowerAlexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch 等NeurIPS 2021 · 被引用 111 次
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 等ICML 2024 · 被引用 44 次
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 被引用 96 次
- Apparently Irrational Choice as Optimal Sequential Decision MakingHaiyang Chen, Hyung Jin Chang, Andrew HowesAAAI 2021 · 被引用 10 次
- An Analytical Study of Utility Functions in Multi-Objective Reinforcement LearningManel Rodriguez-Soto, Juan A. Rodríguez-Aguilar, Maite López-SánchezNeurIPS 2024 · 被引用 9 次
