A Near-Optimal Algorithm for Safe Reinforcement Learning Under Instantaneous Hard Constraints
Ming Shi, Yingbin Liang, Ness B. Shroff
Abstract
In many applications of Reinforcement Learning (RL), it is critically important that the algorithm performs safely, such that instantaneous hard constraints are satisfied at each step, and unsafe states and actions are avoided. However, existing algorithms for "safe" RL are often designed under constraints that either require expected cumulative costs to be bounded or assume all states are safe. Thus, such algorithms could violate instantaneous hard constraints and traverse unsafe states (and actions) in practice. Therefore, in this paper, we develop the first near-optimal safe RL algorithm for episodic Markov Decision Processes with unsafe states and actions under instantaneous hard constraints and the linear mixture model. It not only achieves a regret Õ( dH 3 √ dK ∆c ) that tightly matches the state-of-the-art regret in the setting with only unsafe actions and nearly matches that in the unconstrained setting, but is also safe at each step, where d is the feature-mapping dimension, K is the number of episodes, H is the number of steps in each episode, and ∆ c is a safety-related parameter. We also provide a lower bound Ω(maxdH , which indicates that the dependency on ∆ c is necessary. Further, both our algorithm design and regret analysis involve several novel ideas, which may be of independent interest. Recently, instantaneous hard constraints have been studied in theoretical machine learning. Specifically, [12] and [18] studied bandits with linear instantaneous constraints that require a linear safety value of the chosen action to be bounded at each step. However, it is well-known that bandits are only a very special case of MDP. [14] studied safe linear MDP with linear instantaneous hard constraints. However, they still assume that only the actions could be unsafe, and hence unsafe states (and transitions) are still not considered. Intuitively, when there are only unsafe actions, any action will
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22ab676a-4b33-418a-bf80-03660c86844dCited by top-tier papers6
- Provably Safe Reinforcement Learning with Step-wise Violation ConstraintsNuoya Xiong, Yihan Du, Longbo HuangNeurIPS 2023 · 16 citations
- Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy OptimizationXiyue Peng, Hengquan Guo, Jiawei Zhang, Dongqing Zou et al.NeurIPS 2025 · 9 citations
- Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function ApproximationToshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai et al.NeurIPS 2025 · 5 citations
- Flipping-based Policy for Chance-Constrained Markov Decision ProcessesXun Shen, Shuo Jiang, Akifumi Wachi, Kazumune Hashimoto et al.NeurIPS 2024 · 5 citations
- Primal-Dual Policy Optimization for Linear CMDPs with Adversarial LossesKihyun Yu, Seoungbin Bae, Dabeen LeeICLR 2026 · 2 citations
Builds on12
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
- Provably Efficient Reinforcement Learning for Discounted MDPs with Feature MappingDongruo Zhou, Jiafan He, Quanquan GuICML 2021 · 143 citations
Related papers
- Safe Reinforcement Learning with Linear Function ApproximationSanae Amani, Christos Thrampoulidis, Lin YangICML 2021 · 42 citations
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar et al.NeurIPS 2021 · 110 citations
- Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature SpacesAmirhossein Roknilamouki, Arnob Ghosh, Ming Shi, Fatemeh Nourzad et al.ICML 2025
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningAo Zhou, Jiayi Guan, Li Shen, Fan Lu et al.NeurIPS 2025 · 1 citation
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
