Beyond Average Return in Markov Decision Processes
Alexandre Marthe, Aurélien Garivier, Claire Vernade
Abstract
What are the functionals of the reward that can be computed and optimized exactly in Markov Decision Processes? In the finite-horizon, undiscounted setting, Dynamic Programming (DP) can only handle these operations efficiently for certain classes of statistics. We summarize the characterization of these classes for policy evaluation, and give a new answer for the planning problem. Interestingly, we prove that only generalized means can be optimized exactly, even in the more general framework of Distributional Reinforcement Learning (DistRL). DistRL permits, however, to evaluate other functionals approximately. We provide error bounds on the resulting estimators, and discuss the potential of this approach as well as its limitations. These results contribute to advancing the theory of Markov Decision Processes by examining overall characteristics of the return, and particularly risk-conscious strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 029d6c60-718c-48ef-96b0-c8a32bd9044cCited by top-tier papers6
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 6 citations
- Distributional Bellman Operators over Mean EmbeddingsLi Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter et al.ICML 2024 · 5 citations
- Risk-averse Total-reward MDPs with ERM and EVaRXihong Su, Marek Petrik, Julien Grand-ClémentAAAI 2025 · 3 citations
- Dynamic Programming for Epistemic Uncertainty in Markov Decision ProcessesAxel Benyamine, Julien Grand-Clément, Marek Petrik, Michael Jordan et al.ICML 2026 · 1 citation
- Risk-Averse Total-Reward Reinforcement LearningXihong Su, Jia Lin Hau, Gersi Doko, Kishan Panaganti et al.NeurIPS 2025
Builds on1
Related papers
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 4 citations
- Statistical Efficiency of Distributional Temporal Difference LearningYang Peng, Liangyu Zhang, Zhihua ZhangNeurIPS 2024 · 8 citations
- Distributional Reinforcement Learning via Moment MatchingThanh Nguyen-Tang, Sunil Gupta, Svetha VenkateshAAAI 2021 · 44 citations
- Conjugated Discrete Distributions for Distributional Reinforcement LearningBjörn Lindenberg, Jonas Nordqvist, Karl-Olof LindahlAAAI 2022 · 2 citations
- Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function ApproximationTaehyun Cho, Seungyub Han, Seokhun Ju, Dohyeong Kim et al.ICML 2025
