Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
Mehran Shakerinava, Siamak Ravanbakhsh, Adam M. Oberman
Abstract
Recent work has formalized the reward hypothesis through the lens of expected utility theory, by interpreting reward as utility. Hausner's foundational work showed that dropping the continuity axiom leads to a generalization of expected utility theory where utilities are lexicographically ordered vectors of arbitrary dimension. In this paper, we extend this result by identifying a simple and practical condition under which preferences cannot be represented by scalar rewards, necessitating a 2-dimensional reward function. We provide a full characterization of such reward functions, as well as the general d-dimensional case, in Markov Decision Processes (MDPs) under a memorylessness assumption on preferences. Furthermore, we show that optimal policies in this setting retain many desirable properties of their scalar-reward counterparts, while in the Constrained MDP (CMDP) setting -- another common multiobjective setting -- they do not.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6923725-4368-42db-bb82-06d4c799486fBuilds on3
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 47 citations
- Model-Free Reinforcement Learning for Lexicographic Omega-Regular ObjectivesErnst Moritz Hahn, Mateo Perez, Sven Schewe, Fabio Somenzi et al.FM 2021 · 11 citations
- Utility Theory for Sequential Decision MakingMehran Shakerinava, Siamak RavanbakhshICML 2022 · 8 citations
Related papers
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 21 citations
- Towards Theoretical Understanding of Sequential Decision Making with Preference FeedbackSimone Drago, Marco Mussi, Alberto Maria MetelliICML 2025
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 13 citations
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 6 citations
- An Analytical Study of Utility Functions in Multi-Objective Reinforcement LearningManel Rodriguez-Soto, Juan A. Rodríguez-Aguilar, Maite López-SánchezNeurIPS 2024 · 9 citations
