Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline Policies
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. Ramadge
Abstract
We consider the problem of reinforcement learning when provided with (1) a baseline control policy and (2) a set of constraints that the learner must satisfy. The baseline policy can arise from demonstration data or a teacher agent and may provide useful cues for learning, but it might also be sub-optimal for the task at hand, and is not guaranteed to satisfy the specified constraints, which might encode safety, fairness or other application-specific requirements. In order to safely learn from baseline policies, we propose an iterative policy optimization algorithm that alternates between maximizing expected return on the task, minimizing distance to the baseline policy, and projecting the policy onto the constraint-satisfying set. We analyze our algorithm theoretically and provide a finite-time convergence guarantee. In our experiments on five different control tasks, our algorithm consistently outperforms several state-of-the-art baselines, achieving 10 times fewer constraint violations and 40% higher reward on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2afad945-a387-4845-a483-d6be92fbf98fCited by top-tier papers2
- Grounded Reinforcement Learning: Learning to Win the Game under Human CommandsShusheng Xu, Huaijie Wang, Yi WuNeurIPS 2022 · 5 citations
- Near-optimal Conservative Exploration in Reinforcement Learning under Episode-wise ConstraintsDonghao Li, Ruiquan Huang, Cong Shen, Jing YangICML 2023 · 4 citations
Builds on5
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
- Cautious Adaptation For Reinforcement Learning in Safety-Critical SettingsJesse Zhang, Brian Cheung, Chelsea Finn, Sergey Levine et al.ICML 2020 · 65 citations
- Conservative Safety Critics for ExplorationHomanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine et al.ICLR 2021 · 32 citations
Related papers
- SafeMPO: Constrained Reinforcement Learning with Probabilistic Incremental ImprovementAlexander Mattick, Dominik Seuß, Christopher MutschlerICLR 2026
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPsHarsh Satija, Philip S. Thomas, Joelle Pineau, Romain LarocheNeurIPS 2021 · 30 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
- Constrained Meta Reinforcement Learning with Provable Test-Time SafetyTingting Ni, Maryam KamgarpourICML 2026
