GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic Synthesis
Yushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang, Yan Zheng, Yi Li, Jianye Hao, Yang Liu
摘要
Despite achieving superior performance in human-level control problems, unlike humans, deep reinforcement learning (DRL) lacks high-order intelligence (e.g., logic deduction and reuse), thus it behaves ineffectively than humans regarding learning and generalization in complex problems. Previous works attempt to directly synthesize a white-box logic program as the DRL policy, manifesting logic-driven behaviors. However, most synthesis methods are built on imperative or declarative programming, and each has a distinct limitation, respectively. The former ignores the cause-effect logic during synthesis, resulting in low generalizability across tasks. The latter is strictly proof-based, thus failing to synthesize programs with complex hierarchical logic. In this paper, we combine the above two paradigms together and propose a novel Generalizable Logic Synthesis (GALOIS) framework to synthesize hierarchical and strict cause-effect logic programs. GALOIS leverages the program sketch and defines a new sketch-based hybrid program language for guiding the synthesis. Based on that, GALOIS proposes a sketch-based program synthesis method to automatically generate white-box programs with generalizable and interpretable cause-effect logic. Extensive evaluations on various decision-making tasks with complex logic demonstrate the superiority of GALOIS over mainstream baselines regarding the asymptotic performance, generalizability, and great knowledge reusability across different environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Interpretable and Explainable Logical Policies via Neurally Guided Symbolic AbstractionQuentin Delfosse, Hikaru Shindo, Devendra Singh Dhami, Kristian KerstingNeurIPS 2023 · 被引用 64 次
- Interpretable Concept Bottlenecks to Align Reinforcement Learning AgentsQuentin Delfosse, Sebastian Sztwiertnia, Mark Rothermel, Wolfgang Stammer 等NeurIPS 2024 · 被引用 32 次
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie 等ICLR 2023 · 被引用 5 次
- Logic-Q: Improving Deep Reinforcement Learning-based Quantitative Trading via Program Sketch-based TuningZhiming Li, Junzhe Jiang, Yushi Cao, Aixin Cui 等AAAI 2025 · 被引用 5 次
- Improving Neural Logic Machines via Failure ReflectionZhiming Li, Yushi Cao, Yan Zheng, Xu Liu 等ICML 2024 · 被引用 1 次
它引用的顶会 Paper5
- Program Guided AgentShao-Hua Sun, Te-Lin Wu, Joseph J. LimICLR 2020 · 被引用 63 次
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham 等AAAI 2021 · 被引用 62 次
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon ReasoningDhruv Shah, Peng Xu, Yao Lu, Ted Xiao 等ICLR 2022 · 被引用 50 次
- Provenance-guided synthesis of Datalog programsMukund Raghothaman, Jonathan Mendelson, David Zhao, Mayur Naik 等POPL 2020 · 被引用 49 次
- Program Synthesis Guided Reinforcement Learning for Partially Observed EnvironmentsYichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu 等NeurIPS 2021
相关 Paper
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- Integrating Planning and Deep Reinforcement Learning via Automatic Induction of Task SubstructuresJung-Chun Liu, Chi-Hsien Chang, Shao-Hua Sun, Tian-Li YuICLR 2024 · 被引用 6 次
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 被引用 42 次
- Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement LearningHengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang 等ACL 2026
- In a Nutshell, the Human Asked for This: Latent Goals for Following Temporal SpecificationsBorja G. León, Murray Shanahan, Francesco BelardinelliICLR 2022 · 被引用 23 次
