Constrained Markov Decision Processes via Backward Value Functions
Harsh Satija, Philip Amortila, Joelle Pineau
Abstract
Although Reinforcement Learning (RL) algorithms have found tremendous success in simulated domains, they often cannot directly be applied to physical systems, especially in cases where there are hard constraints to satisfy (e.g. on safety or resources). In standard RL, the agent is incentivized to explore any behavior as long as it maximizes rewards, but in the real world, undesired behavior can damage either the system or the agent in a way that breaks the learning process itself. In this work, we model the problem of learning with constraints as a Constrained Markov Decision Process and provide a new on-policy formulation for solving it. A key contribution of our approach is to translate cumulative cost constraints into state-based constraints. Through this, we define a safe policy improvement method which maximizes returns while ensuring that the constraints are satisfied at every step. We provide theoretical guarantees under which the agent converges while ensuring safety over the course of training. We also highlight the computational advantages of this approach. The effectiveness of our approach is demonstrated on safe navigation tasks and in safety-constrained versions of Mu-JoCo environments, with deep neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28185575-1005-4c99-a548-05b4bd37f35dCited by top-tier papers23
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang et al.NeurIPS 2022 · 95 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
- Learning with Safety Constraints: Sample Complexity of Reinforcement Learning for Constrained MDPsAria HasanzadeZonuzy, Archana Bura, Dileep M. Kalathil, Srinivas ShakkottaiAAAI 2021 · 46 citations
- Forethought and Hindsight in Credit AssignmentVeronica Chelu, Doina Precup, Hado van HasseltNeurIPS 2020 · 29 citations
Builds on1
Related papers
- SafeMPO: Constrained Reinforcement Learning with Probabilistic Incremental ImprovementAlexander Mattick, Dominik Seuß, Christopher MutschlerICLR 2026
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- Conservative and Adaptive Penalty for Model-Based Safe Reinforcement LearningYecheng Jason Ma, Andrew Shen, Osbert Bastani, Dinesh JayaramanAAAI 2022 · 32 citations
- Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization AlgorithmAshish Kumar Jayant, Shalabh BhatnagarNeurIPS 2022 · 84 citations
- Balancing Constraints and Rewards with Meta-Gradient D4PGDan A. Calian, Daniel J. Mankowitz, Tom Zahavy, Zhongwen Xu et al.ICLR 2021 · 6 citations
