Reachability Constrained Reinforcement Learning
Dongjie Yu, Haitong Ma, Sheng-bo Li, Jianyu Chen
Abstract
Constrained reinforcement learning (CRL) has gained significant interest recently, since safety constraints satisfaction is critical for real-world problems. However, existing CRL methods constraining discounted cumulative costs generally lack rigorous definition and guarantee of safety. In contrast, in the safe control research, safety is defined as persistently satisfying certain state constraints. Such persistent safety is possible only on a subset of the state space, called feasible set, where an optimal largest feasible set exists for a given environment. Recent studies incorporate feasible sets into CRL with energy-based methods such as control barrier function (CBF), safety index (SI), and leverage prior conservative estimations of feasible sets, which harms the performance of the learned policy. To deal with this problem, this paper proposes the reachability CRL (RCRL) method by using reachability analysis to establish the novel self-consistency condition and characterize the feasible sets. The feasible sets are represented by the safety value function, which is used as the constraint in CRL. We use the multi-time scale stochastic approximation theory to prove that the proposed algorithm converges to a local optimum, where the largest feasible set can be guaranteed. Empirical results on different benchmarks validate the learned feasible set, the policy performance, and constraint satisfaction of RCRL, compared to CRL and safe control baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65efc877-3941-4a28-9d7e-e74b44509dbdCited by top-tier papers32
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang et al.ICML 2023 · 81 citations
- Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion ModelYinan Zheng, Jianxiong Li, Dongjie Yu, Yujie Yang et al.ICLR 2024 · 72 citations
- Iterative Reachability Estimation for Safe Reinforcement LearningMilan Ganai, Zheng Gong, Chenning Yu, Sylvia L. Herbert et al.NeurIPS 2023 · 56 citations
- Solving Minimum-Cost Reach Avoid using Reinforcement LearningOswin So, Cheng Ge, Chuchu FanNeurIPS 2024 · 24 citations
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li et al.ICML 2023 · 19 citations
Builds on7
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
- Safe Reinforcement Learning by Imagining the Near FutureGarrett Thomas, Yuping Luo, Tengyu MaNeurIPS 2021 · 118 citations
- Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time ViolationsYuping Luo, Tengyu MaNeurIPS 2021 · 58 citations
Related papers
- A Physics-Informed Machine Learning Framework for Safe and Optimal Control of Autonomous SystemsManan Tayal, Aditya Singh, Shishir Kolathaya, Somil BansalICML 2025
- SafeMPO: Constrained Reinforcement Learning with Probabilistic Incremental ImprovementAlexander Mattick, Dominik Seuß, Christopher MutschlerICLR 2026
- Feasibility Consistent Representation Learning for Safe Reinforcement LearningZhepeng Cen, Yihang Yao, Zuxin Liu, Ding ZhaoICML 2024 · 3 citations
- ActSafe: Active Exploration with Safety Constraints for Reinforcement LearningYarden As, Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza et al.ICLR 2025
- Feasible Reachable Policy IterationShentao Qin, Yujie Yang, Yao Mu, Jie Li et al.ICML 2024 · 3 citations
