Planning and Learning in Average Risk-aware MDPs
Weikai Wang, Erick Delage
Abstract
For continuing tasks, average cost Markov decision processes have well-documented value and can be solved using efficient algorithms. However, it explicitly assumes that the agent is risk-neutral. In this work, we extend risk-neutral algorithms to accommodate the more general class of dynamic risk measures. Specifically, we propose a relative value iteration (RVI) algorithm for planning and design two model-free Q-learning algorithms, namely a generic algorithm based on the multi-level Monte Carlo (MLMC) method, and an off-policy algorithm dedicated to utility-based shortfall risk measures. Both the RVI and MLMC-based Q-learning algorithms are proven to converge to optimality. Numerical experiments validate our analysis, confirm empirically the convergence of the off-policy algorithm, and demonstrate that our approach enables the identification of policies that are finely tuned to the intricate risk-awareness of the agent that they serve.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e705fe33-0784-48b6-8aa6-f15403dfd33aBuilds on3
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong et al.ICML 2022 · 72 citations
- Risk-Aware Reinforcement Learning with Coherent Risk Measures and Non-linear Function ApproximationThanh Lam, Arun Verma, Bryan Kian Hsiang Low, Patrick JailletICLR 2023
Related papers
- Reinforcement Learning for Cost-Aware Markov Decision ProcessesWesley Suttle, Kaiqing Zhang, Zhuoran Yang, Ji Liu et al.ICML 2021 · 11 citations
- Risk-Averse Total-Reward Reinforcement LearningXihong Su, Jia Lin Hau, Gersi Doko, Kishan Panaganti et al.NeurIPS 2025
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang et al.NeurIPS 2020 · 87 citations
- Model-Free Robust Average-Reward Reinforcement LearningYue Wang, Alvaro Velasquez, George K. Atia, Ashley Prater-Bennette et al.ICML 2023 · 25 citations
- Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinityAneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon et al.ICML 2026 · 1 citation
