Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline Policies
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. Ramadge
摘要
We consider the problem of reinforcement learning when provided with (1) a baseline control policy and (2) a set of constraints that the learner must satisfy. The baseline policy can arise from demonstration data or a teacher agent and may provide useful cues for learning, but it might also be sub-optimal for the task at hand, and is not guaranteed to satisfy the specified constraints, which might encode safety, fairness or other application-specific requirements. In order to safely learn from baseline policies, we propose an iterative policy optimization algorithm that alternates between maximizing expected return on the task, minimizing distance to the baseline policy, and projecting the policy onto the constraint-satisfying set. We analyze our algorithm theoretically and provide a finite-time convergence guarantee. In our experiments on five different control tasks, our algorithm consistently outperforms several state-of-the-art baselines, achieving 10 times fewer constraint violations and 40% higher reward on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Grounded Reinforcement Learning: Learning to Win the Game under Human CommandsShusheng Xu, Huaijie Wang, Yi WuNeurIPS 2022 · 被引用 5 次
- Near-optimal Conservative Exploration in Reinforcement Learning under Episode-wise ConstraintsDonghao Li, Ruiquan Huang, Cong Shen, Jing YangICML 2023 · 被引用 4 次
它引用的顶会 Paper5
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 被引用 190 次
- Cautious Adaptation For Reinforcement Learning in Safety-Critical SettingsJesse Zhang, Brian Cheung, Chelsea Finn, Sergey Levine 等ICML 2020 · 被引用 65 次
- Conservative Safety Critics for ExplorationHomanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine 等ICLR 2021 · 被引用 32 次
相关 Paper
- SafeMPO: Constrained Reinforcement Learning with Probabilistic Incremental ImprovementAlexander Mattick, Dominik Seuß, Christopher MutschlerICLR 2026
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPsHarsh Satija, Philip S. Thomas, Joelle Pineau, Romain LarocheNeurIPS 2021 · 被引用 30 次
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- Constrained Meta Reinforcement Learning with Provable Test-Time SafetyTingting Ni, Maryam KamgarpourICML 2026
