Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
Mehran Shakerinava, Siamak Ravanbakhsh, Adam M. Oberman
摘要
Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 被引用 47 次
- Model-Free Reinforcement Learning for Lexicographic Omega-Regular ObjectivesErnst Moritz Hahn, Mateo Perez, Sven Schewe, Fabio Somenzi 等FM 2021 · 被引用 11 次
- Utility Theory for Sequential Decision MakingMehran Shakerinava, Siamak RavanbakhshICML 2022 · 被引用 8 次
相关 Paper
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 被引用 21 次
- Towards Theoretical Understanding of Sequential Decision Making with Preference FeedbackSimone Drago, Marco Mussi, Alberto Maria MetelliICML 2025
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 被引用 13 次
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 被引用 6 次
- An Analytical Study of Utility Functions in Multi-Objective Reinforcement LearningManel Rodriguez-Soto, Juan A. Rodríguez-Aguilar, Maite López-SánchezNeurIPS 2024 · 被引用 9 次
