Programmatic Reinforcement Learning without Oracles
Wenjie Qiu, He Zhu
摘要
Deep reinforcement learning (RL) has led to encouraging successes in many challenging control tasks. However, a deep RL model lacks interpretability due to the difficulty of identifying how the model's control logic relates to its network structure. Programmatic policies structured in more interpretable representations emerge as a promising solution. Yet two shortcomings remain: First, synthesizing programmatic policies requires optimizing over the discrete and non-differentiable search space of program architectures. Previous works are suboptimal because they only enumerate program architectures greedily guided by a pretrained RL oracle. Second, these works do not exploit compositionality, an important programming concept, to reuse and compose primitive functions to form a complex function for new tasks. Our first contribution is a programmatically interpretable RL framework that conducts program architecture search on top of a continuous relaxation of the architecture space defined by programming language grammar rules. Our algorithm allows policy architectures to be learned with policy parameters via bilevel optimization using efficient policy-gradient methods, and thus does not require a pretrained oracle. Our second contribution is improving programmatic policies to support compositionality by integrating primitive functions learned to grasp task-agnostic skills as a composite program to solve novel RL problems. Experiment results demonstrate that our algorithm excels in discovering optimal programmatic policies that are highly interpretable.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper16
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree SearchNicola Dainese, Matteo Merler, Minttu Alakuijala, Pekka MarttinenNeurIPS 2024 · 被引用 49 次
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li 等NeurIPS 2024 · 被引用 48 次
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 被引用 44 次
- End-to-End Neuro-Symbolic Reinforcement Learning with Textual ExplanationsLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang 等ICML 2024 · 被引用 18 次
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi 等NeurIPS 2023 · 被引用 15 次
相关 Paper
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- Learning Differentiable Programs with Admissible Neural HeuristicsAmeesh Shah, Eric Zhan, Jennifer J. Sun, Abhinav Verma 等NeurIPS 2020 · 被引用 56 次
- Hierarchical Programmatic Option FrameworkYu-An Lin, Chen-Tao Lee, Chih-Han Yang, Guan-Ting Liu 等NeurIPS 2024 · 被引用 7 次
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson 等ICLR 2020 · 被引用 35 次
- Differentiable Synthesis of Program ArchitecturesGuofeng Cui, He ZhuNeurIPS 2021 · 被引用 20 次
