Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability
Whiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul Sung
Abstract
Constrained reinforcement learning (RL) is an area of RL whose objective is to find an optimal policy that maximizes expected cumulative return while satisfying a given constraint. Most of the previous constrained RL works consider expected cumulative sum cost as the constraint. However, optimization with this constraint cannot guarantee a target probability of outage event that the cumulative sum cost exceeds a given threshold. This paper proposes a framework, named Quantile Constrained RL (QCRL), to constrain the quantile of the distribution of the cumulative sum cost that is a necessary and sufficient condition to satisfy the outage constraint. This is the first work that tackles the issue of applying the policy gradient theorem to the quantile and provides theoretical results for approximating the gradient of the quantile. Based on the derived theoretical results and the technique of the Lagrange multiplier, we construct a constrained RL algorithm named Quantile Constrained Policy Optimization (QCPO). We use distributional RL with the Large Deviation Principle (LDP) to estimate quantiles and tail probability of the cumulative sum cost for the implementation of QCPO. The implemented algorithm satisfies the outage probability constraint after the training period.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd585401-a212-4dc2-bc44-682672c0a431Cited by top-tier papers4
- Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble AgentsWoojun Kim, Yongjae Shin, Jongeui Park, Youngchul SungNeurIPS 2023 · 20 citations
- Constrained Multi-Objective Reinforcement Learning with Max-Min CriterionGiseung Park, Hyunyoung Nam, Woohyeon Byeon, Amir Leshem et al.ICML 2026
- Online Learning in Risk Sensitive constrained MDPArnob Ghosh, Mehrdad MoharramiICML 2025
- Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement LearningChenglin Li, Guangchun Ruan, Hua GengAAAI 2025
Builds on8
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- IPO: Interior-Point Policy Optimization under ConstraintsYongshuai Liu, Jiaxin Ding, Xin LiuAAAI 2020 · 231 citations
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
Related papers
- Extreme Value Policy Optimization for Safe Reinforcement LearningShiqing Gao, Yihang Zhou, Shuai Shao, Haoyu Luo et al.ICML 2025
- Off-Policy Safe Reinforcement Learning with Cost-Constrained Optimistic ExplorationGuopeng Li, Matthijs T. J. Spaan, Julian F. P. KooijICLR 2026
- POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement LearningJiayi Guan, Li Shen, Ao Zhou, Lusong Li et al.CVPR 2024
- Risk-Averse Constrained Reinforcement Learning with Optimized Certainty EquivalentsJane H. Lee, Baturay Saglam, Spyridon Pougkakiotis, Amin Karbasi et al.NeurIPS 2025 · 2 citations
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
