Guiding Safe Exploration with Weakest Preconditions
Greg Anderson, Swarat Chaudhuri, Isil Dillig
摘要
In reinforcement learning for safety-critical settings, it is often desirable for the agent to obey safety constraints at all points in time, including during training. We present a novel neurosymbolic approach called SPICE to solve this safe exploration problem. SPICE uses an online shielding layer based on symbolic weakest preconditions to achieve a more precise safety analysis than existing tools without unduly impacting the training process. We evaluate the approach on a suite of continuous control benchmarks and show that it can achieve comparable performance to existing safe learning techniques while incurring fewer safety violations. Additionally, we present theoretical results showing that SPICE converges to the optimal safe policy under reasonable assumptions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Dynamic Model Predictive Shielding for Provably Safe Reinforcement LearningArko Banerjee, Kia Rahmani, Joydeep Biswas, Isil DilligNeurIPS 2024 · 被引用 24 次
- AED: Adaptable Error Detection for Few-shot Imitation PolicyJia-Fong Yeh, Kuo-Han Hung, Pang-Chi Lo, Chi-Ming Chung 等NeurIPS 2024 · 被引用 3 次
- Safe Exploration in Reinforcement Learning by Reachability Analysis over Learned ModelsYuning Wang, He ZhuCAV 2024 · 被引用 2 次
- Robust Adaptive Multi-Step Predictive ShieldingTanmay Ambadkar, Darshan Chudiwal, Greg Anderson, Abhinav VermaICLR 2026
它引用的顶会 Paper8
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 被引用 238 次
- WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement LearningQisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. SpaanAAAI 2021 · 被引用 168 次
- Neurosymbolic Reinforcement Learning with Formally Verified ExplorationGreg Anderson, Abhinav Verma, Isil Dillig, Swarat ChaudhuriNeurIPS 2020 · 被引用 91 次
- Reachability Constrained Reinforcement LearningDongjie Yu, Haitong Ma, Sheng-bo Li, Jianyu ChenICML 2022 · 被引用 90 次
相关 Paper
- Safe Neurosymbolic Learning with Differentiable Symbolic ExecutionChenxi Yang, Swarat ChaudhuriICLR 2022 · 被引用 13 次
- Enhancing Safe Exploration Using Safety State AugmentationAivar Sootla, Alexander I. Cowen-Rivers, Jun Wang, Haitham Bou-AmmarNeurIPS 2022 · 被引用 23 次
- Safety Representations for Safer Policy LearningKaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen 等ICLR 2025
- Probabilistic Shielding for Safe Reinforcement LearningEdwin Hamel-De le Court, Francesco Belardinelli, Alexander W. GoodallAAAI 2025 · 被引用 7 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
