Neurosymbolic Reinforcement Learning with Formally Verified Exploration
Greg Anderson, Abhinav Verma, Isil Dillig, Swarat Chaudhuri
Abstract
We present Revel, a partially neural reinforcement learning (RL) framework for provably safe exploration in continuous state and action spaces. A key challenge for provably safe deep RL is that repeatedly verifying neural networks within a learning loop is computationally infeasible. We address this challenge using two policy classes: a general, neurosymbolic class with approximate gradients and a more restricted class of symbolic policies that allows efficient verification. Our learning algorithm is a mirror descent over policies: in each iteration, it safely lifts a symbolic policy into the neurosymbolic space, performs safe gradient updates to the resulting policy, and projects the updated policy into the safe symbolic subset, all without requiring explicit verification of neural networks. Our empirical results show that Revel enforces safe exploration in many scenarios in which Constrained Policy Optimization does not, and that it can discover policies that outperform those learned through prior approaches to verified exploration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2e20451-381e-4edc-a5a7-c45a1c8a2f4dCited by top-tier papers16
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 112 citations
- Safe Reinforcement Learning Using Advantage-Based InterventionNolan Wagener, Byron Boots, Ching-An ChengICML 2021 · 66 citations
- Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time ViolationsYuping Luo, Tengyu MaNeurIPS 2021 · 58 citations
- Compositional Policy Learning in Stochastic Control Systems with Formal GuaranteesDorde Zikelic, Mathias Lechner, Abhinav Verma, Krishnendu Chatterjee et al.NeurIPS 2023 · 31 citations
- Web question answering with neurosymbolic program synthesisQiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett et al.PLDI 2021 · 25 citations
Builds on2
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang et al.USENIX Security 2018 · 523 citations
Related papers
- Guiding Safe Exploration with Weakest PreconditionsGreg Anderson, Swarat Chaudhuri, Isil DilligICLR 2023
- Safe Exploration in Reinforcement Learning by Reachability Analysis over Learned ModelsYuning Wang, He ZhuCAV 2024 · 2 citations
- Discovering symbolic policies with deep reinforcement learningMikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago et al.ICML 2021 · 118 citations
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi et al.NeurIPS 2023 · 15 citations
- Safe DNN-type Controller Synthesis for Nonlinear Systems via Meta Reinforcement LearningHanrui Zhao, Xia Zeng, Niuniu Qi, Zhengfeng Yang et al.DAC 2023 · 4 citations
