Reinforcement Learning of Risk-Constrained Policies in Markov Decision Processes
Tomás Brázdil, Krishnendu Chatterjee, Petr Novotný, Jiri Vahala
摘要
Markov decision processes (MDPs) are the defacto framework for sequential decision making in the presence of stochastic uncertainty. A classical optimization criterion for MDPs is to maximize the expected discounted-sum payoff, which ignores low probability catastrophic events with highly negative impact on the system. On the other hand, risk-averse policies require the probability of undesirable events to be below a given threshold, but they do not account for optimization of the expected payoff. We consider MDPs with discounted-sum payoff with failure states which represent catastrophic outcomes. The objective of risk-constrained planning is to maximize the expected discounted-sum payoff among risk-averse policies that ensure the probability to encounter a failure state is below a desired threshold. Our main contribution is an efficient risk-constrained planning algorithm that combines UCT-like search with a predictor learned through interaction with the MDP (in the style of AlphaZero) and with a risk-constrained action selection via linear programming. We demonstrate the effectiveness of our approach with experiments on classical MDPs from the literature, including benchmarks with an order of 10 6 states.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Threshold UCT: Cost-Constrained Monte Carlo Tree Search with Pareto CurvesMartin Kurecka, Václav Nevyhostený, Petr Novotný, Vít UncovskýAAAI 2025 · 被引用 1 次
- Shields to Guarantee Probabilistic Safety in MDPsLinus Heck, Filip Macák, Roman Andriushchenko, Milan Ceska 等CAV 2026
相关 Paper
- RiskZero: Plan More to Risk Less with a Learned ModelYousef Yassin, Junfeng WenICML 2026
- A Sample-Efficient Algorithm for Episodic Finite-Horizon MDP with ConstraintsKrishna Chaitanya Kalagarla, Rahul Jain, Pierluigi NuzzoAAAI 2021 · 被引用 58 次
- Risk-Averse Constrained Reinforcement Learning with Optimized Certainty EquivalentsJane H. Lee, Baturay Saglam, Spyridon Pougkakiotis, Amin Karbasi 等NeurIPS 2025 · 被引用 2 次
- PAC Statistical Model Checking of Mean Payoff in Discrete- and Continuous-Time MDPChaitanya Agarwal, Shibashis Guha, Jan Kretínský, Pazhamalai MuruganandhamCAV 2022 · 被引用 8 次
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert 等ICML 2020 · 被引用 84 次
