Hierarchical Programmatic Option Framework
Yu-An Lin, Chen-Tao Lee, Chih-Han Yang, Guan-Ting Liu, Shao-Hua Sun
Abstract
Deep reinforcement learning aims to learn deep neural network policies to solve large-scale decision-making problems. However, approximating policies using deep neural networks makes it difficult to interpret the learned decision-making process. To address this issue, prior works [10, 46, 74] proposed to use human-readable programs as policies to increase the interpretability of the decision-making pipeline. Nevertheless, programmatic policies generated by these methods struggle to effectively solve long and repetitive RL tasks and cannot generalize to even longer horizons during testing. To solve these problems, we propose the Hierarchical Programmatic Option framework (HIPO), which aims to solve long and repetitive RL problems with human-readable programs as options (low-level policies). Specifically, we propose a method that retrieves a set of effective, diverse, and compatible programs as options. Then, we learn a high-level policy to effectively reuse these programmatic options to solve reoccurring subtasks. Our proposed framework outperforms programmatic RL and deep RL baselines on various tasks. Ablation studies justify the effectiveness of our proposed search algorithm for retrieving a set of programmatic options.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7baea0a-caa3-4530-9609-22ccc7287050Cited by top-tier papers2
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to DeliberationZhiwei Zhang, Xiaomin Li, Yudi Lin, Hui Liu et al.ICLR 2026 · 13 citations
- Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided SearchMax Liu, Chan-Hung Yu, Wei-Hsu Lee, Cheng-Wei Hung et al.ICLR 2025
Builds on21
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Discovering symbolic policies with deep reinforcement learningMikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago et al.ICML 2021 · 118 citations
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 104 citations
- Learning Robot Skills with Temporal Variational InferenceTanmay Shankar, Abhinav GuptaICML 2020 · 80 citations
- Demo2Code: From Summarizing Demonstrations to Synthesizing Code via Extended Chain-of-ThoughtYuki Wang, Gonzalo Gonzalez-Pumariega, Yash Sharma, Sanjiban ChoudhuryNeurIPS 2023 · 67 citations
Related papers
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 42 citations
- Hierarchical Programmatic Reinforcement Learning via Learning to Compose ProgramsGuan-Ting Liu, En-Pei Hu, Pu-Jen Cheng, Hung-Yi Lee et al.ICML 2023 · 21 citations
- Creativity of AI: Automatic Symbolic Option Discovery for Facilitating Deep Reinforcement LearningMu Jin, Zhihao Ma, Kebing Jin, Hankz Hankui Zhuo et al.AAAI 2022 · 49 citations
- Offline Hierarchical Reinforcement Learning via Inverse OptimizationCarolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone et al.ICLR 2025
- GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisYushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang et al.NeurIPS 2022 · 23 citations
