Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards
Silviu Pitis
Abstract
As the capabilities of artificial agents improve, they are being increasingly deployed to service multiple diverse objectives and stakeholders. However, the composition of these objectives is often performed ad hoc, with no clear justification. This paper takes a normative approach to multi-objective agency: from a set of intuitively appealing axioms, it is shown that Markovian aggregation of Markovian reward functions is not possible when the time preference (discount factor) for each objective may vary. It follows that optimal multi-objective agents must admit rewards that are non-Markovian with respect to the individual objectives. To this end, a practical non-Markovian aggregation scheme is proposed, which overcomes the impossibility with only one additional parameter for each objective. This work offers new insights into sequential, multi-objective agency and intertemporal choice, and has practical implications for the design of AI systems deployed to serve multiple generations of principals with varying time preference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8e4b3dd-eaa0-486b-b785-6e45a131a894Cited by top-tier papers2
- Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman FormulationsWilliam Sharpless, Dylan Hirsch, Sander Tonkens, Nikhil Uday Shinde et al.ICLR 2026 · 12 citations
- A Black Swan Hypothesis: The Role of Human Irrationality in AI SafetyHyunin Lee, Chanwoo Park, David Abel, Ming JinICLR 2025
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fine-tuning language models to find agreement among humans with diverse preferencesMichiel A. Bakker, Martin J. Chadwick, Hannah Sheahan, Michael Henry Tessler et al.NeurIPS 2022 · 349 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Compositional Visual Generation with Energy Based ModelsYilun Du, Shuang Li, Igor MordatchNeurIPS 2020 · 225 citations
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
Related papers
- Utility Theory for Sequential Decision MakingMehran Shakerinava, Siamak RavanbakhshICML 2022 · 8 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- A Continuous-time Tractable Model for Present-biased AgentsYasunori Akagi, Hideaki Kim, Takeshi KurashimaAAAI 2025 · 2 citations
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 6 citations
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm et al.ICLR 2024 · 3 citations
