Risk-averse Total-reward MDPs with ERM and EVaR
Xihong Su, Marek Petrik, Julien Grand-Clément
摘要
Optimizing risk-averse objectives in discounted MDPs is challenging because most models do not admit direct dynamic programming equations and require complex history-dependent policies. In this paper, we show that the risk-averse total reward criterion, under the Entropic Risk Measure (ERM) and Entropic Value at Risk (EVaR) risk measures, can be optimized by a stationary policy, making it simple to analyze, interpret, and deploy. We propose exponential value iteration, policy iteration, and linear programming to compute optimal policies. Compared with prior work, our results only require the relatively mild condition of transient MDPs and allow for both positive and negative rewards. Our results indicate that the total reward criterion may be preferable to the discounted criterion in a broad range of risk-averse reinforcement learning domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 被引用 70 次
- Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing ApproachYingjie Fei, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 53 次
- Minimax Regret for Stochastic Shortest PathAlon Cohen, Yonathan Efroni, Yishay Mansour, Aviv RosenbergNeurIPS 2021 · 被引用 32 次
- Constrained Risk-Averse Markov Decision ProcessesMohamadreza Ahmadi, Ugo Rosolia, Michel D. Ingham, Richard M. Murray 等AAAI 2021 · 被引用 31 次
- On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision ProcessesJia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek PetrikNeurIPS 2023 · 被引用 22 次
相关 Paper
- Dynamic Programming for Epistemic Uncertainty in Markov Decision ProcessesAxel Benyamine, Julien Grand-Clément, Marek Petrik, Michael Jordan 等ICML 2026 · 被引用 1 次
- Risk-Aware Stochastic Shortest PathTobias MeggendorferAAAI 2022 · 被引用 13 次
- On the Global Convergence of Risk-Averse Policy Gradient Methods with Expected Conditional Risk MeasuresXian Yu, Lei YingICML 2023 · 被引用 8 次
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 被引用 44 次
- Risk-Sensitive Variational Actor-Critic: A Model-Based ApproachAlonso Granados Baca, Reza Ebrahimi, Jason PachecoICLR 2025
