GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic Synthesis
Yushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang, Yan Zheng, Yi Li, Jianye Hao, Yang Liu
Abstract
Despite achieving superior performance in human-level control problems, unlike humans, deep reinforcement learning (DRL) lacks high-order intelligence (e.g., logic deduction and reuse), thus it behaves ineffectively than humans regarding learning and generalization in complex problems. Previous works attempt to directly synthesize a white-box logic program as the DRL policy, manifesting logic-driven behaviors. However, most synthesis methods are built on imperative or declarative programming, and each has a distinct limitation, respectively. The former ignores the cause-effect logic during synthesis, resulting in low generalizability across tasks. The latter is strictly proof-based, thus failing to synthesize programs with complex hierarchical logic. In this paper, we combine the above two paradigms together and propose a novel Generalizable Logic Synthesis (GALOIS) framework to synthesize hierarchical and strict cause-effect logic programs. GALOIS leverages the program sketch and defines a new sketch-based hybrid program language for guiding the synthesis. Based on that, GALOIS proposes a sketch-based program synthesis method to automatically generate white-box programs with generalizable and interpretable cause-effect logic. Extensive evaluations on various decision-making tasks with complex logic demonstrate the superiority of GALOIS over mainstream baselines regarding the asymptotic performance, generalizability, and great knowledge reusability across different environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Interpretable and Explainable Logical Policies via Neurally Guided Symbolic AbstractionQuentin Delfosse, Hikaru Shindo, Devendra Singh Dhami, Kristian KerstingNeurIPS 2023 · 64 citations
- Interpretable Concept Bottlenecks to Align Reinforcement Learning AgentsQuentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer et al.NeurIPS 2024 · 32 citations
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie et al.ICLR 2023 · 5 citations
- Logic-Q: Improving Deep Reinforcement Learning-based Quantitative Trading via Program Sketch-based TuningZhiming Li, Junzhe Jiang, Yushi Cao, Aixin Cui et al.AAAI 2025 · 5 citations
- Improving Neural Logic Machines via Failure ReflectionZhiming Li, Yushi Cao, Yan Zheng, Xu Liu et al.ICML 2024 · 1 citation
Builds on5
- Program Guided AgentShao-Hua Sun, Te-Lin Wu, Joseph J. LimICLR 2020 · 63 citations
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham et al.AAAI 2021 · 62 citations
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon ReasoningDhruv Shah, Peng Xu, Yao Lu, Ted Xiao et al.ICLR 2022 · 50 citations
- Provenance-guided synthesis of Datalog programsMukund Raghothaman, Jonathan Mendelson, David Zhao, Mayur Naik et al.POPL 2020 · 49 citations
- Program Synthesis Guided Reinforcement Learning for Partially Observed EnvironmentsYichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu et al.NeurIPS 2021
Related papers
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 104 citations
- Integrating Planning and Deep Reinforcement Learning via Automatic Induction of Task SubstructuresJung-Chun Liu, Chi-Hsien Chang, Shao-Hua Sun, Tian-Li YuICLR 2024 · 6 citations
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 42 citations
- Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement LearningHengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang et al.ACL 2026
- In a Nutshell, the Human Asked for This: Latent Goals for Following Temporal SpecificationsBorja G. León, Murray Shanahan, Francesco BelardinelliICLR 2022 · 23 citations
