Hierarchical Programmatic Reinforcement Learning via Learning to Compose Programs
Guan-Ting Liu, En-Pei Hu, Pu-Jen Cheng, Hung-Yi Lee, Shao-Hua Sun
摘要
Aiming to produce reinforcement learning (RL) policies that are human-interpretable and can generalize better to novel scenarios, Trivedi et al. ( 2021 ) present a method (LEAPS) that first learns a program embedding space to continuously parameterize diverse programs from a pre-generated program dataset, and then searches for a tasksolving program in the learned program embedding space when given a task. Despite the encouraging results, the program policies that LEAPS can produce are limited by the distribution of the program dataset. Furthermore, during searching, LEAPS evaluates each candidate program solely based on its return, failing to precisely reward correct parts of programs and penalize incorrect parts. To address these issues, we propose to learn a meta-policy that composes a series of programs sampled from the learned program embedding space. By learning to compose programs, our proposed hierarchical programmatic reinforcement learning (HPRL) framework can produce program policies that describe out-of-distributionally complex behaviors and directly assign credits to programs that induce desired behaviors. The experimental results in the Karel domain show that our proposed framework outperforms baselines. The ablation studies confirm the limitations of LEAPS and justify our design choices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree SearchNicola Dainese, Matteo Merler, Minttu Alakuijala, Pekka MarttinenNeurIPS 2024 · 被引用 49 次
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li 等NeurIPS 2024 · 被引用 48 次
- Reclaiming the Source of Programmatic Policies: Programmatic versus Latent SpacesTales Henrique Carvalho, Kenneth Tjhia, Levi LelisICLR 2024 · 被引用 8 次
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control PoliciesQinglong Hu, Tong Xialiang, Mingxuan Yuan, Fei Liu 等ICLR 2026 · 被引用 7 次
- Hierarchical Programmatic Option FrameworkYu-An Lin, Chen-Tao Lee, Chih-Han Yang, Guan-Ting Liu 等NeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper11
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- TreeGen: A Tree-Based Transformer Architecture for Code GenerationZeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun 等AAAI 2020 · 被引用 196 次
- Jigsaw: Large Language Models meet Program SynthesisNaman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan 等ICSE 2022 · 被引用 134 次
- Discovering symbolic policies with deep reinforcement learningMikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago 等ICML 2021 · 被引用 118 次
相关 Paper
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 被引用 42 次
- Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided SearchMax Liu, Chan-Hung Yu, Wei-Hsu Lee, Cheng-Wei Hung 等ICLR 2025
- Revisiting OOD Generalization in Programmatic RLAmirhossein Rajabpour, Kiarash Aghakasiri, Sandra Zilles, Levi LelisICML 2026
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi 等NeurIPS 2023 · 被引用 15 次
