Exploring Safer Behaviors for Deep Reinforcement Learning
Enrico Marchesini, Davide Corsi, Alessandro Farinelli
摘要
We consider Reinforcement Learning (RL) problems where an agent attempts to maximize a reward signal while minimizing a cost function that models unsafe behaviors. Such formalization is addressed in the literature using constrained optimization on the cost, limiting the exploration and leading to a significant trade-off between cost and reward. In contrast, we propose a Safety-Oriented Search that complements Deep RL algorithms to bias the policy toward safety within an evolutionary cost optimization. We leverage evolutionary exploration benefits to design a novel concept of safe mutations that use visited unsafe states to explore safer actions. We further characterize the behaviors of the policies over desired specifics with a sample-based bound estimation, which makes prior verification analysis tractable in the training loop. Hence, driving the learning process towards safer regions of the policy space. Empirical evidence on the Safety Gym benchmark shows that we successfully avoid drawbacks on the return while improving the safety of the policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Direct Behavior Specification via Constrained Reinforcement LearningJulien Roy, Roger Girgis, Joshua Romoff, Pierre-Luc Bacon 等ICML 2022 · 被引用 46 次
- Improving Deep Policy Gradients with Value Function SearchEnrico Marchesini, Christopher AmatoICLR 2023
它引用的顶会 Paper4
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang 等USENIX Security 2018 · 被引用 523 次
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- Genetic Soft Updates for Policy Evolution in Deep Reinforcement LearningEnrico Marchesini, Davide Corsi, Alessandro FarinelliICLR 2021 · 被引用 33 次
相关 Paper
- Safe Exploration in Reinforcement Learning: A Generalized Formulation and AlgorithmsAkifumi Wachi, Wataru Hashimoto, Xun Shen, Kazumune HashimotoNeurIPS 2023 · 被引用 38 次
- Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization AlgorithmAshish Kumar Jayant, Shalabh BhatnagarNeurIPS 2022 · 被引用 84 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
- Enhancing Efficiency of Safe Reinforcement Learning via Sample ManipulationShangding Gu, Laixi Shi, Yuhao Ding, Alois Knoll 等NeurIPS 2024 · 被引用 14 次
- Safety Representations for Safer Policy LearningKaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen 等ICLR 2025
