Sub-policy Adaptation for Hierarchical Reinforcement Learning
Alexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter Abbeel
Abstract
Hierarchical reinforcement learning is a promising approach to tackle long-horizon decision-making problems with sparse rewards. Unfortunately, most methods still decouple the lower-level skill acquisition process and the training of a higher level that controls the skills in a new task. Leaving the skills fixed can lead to significant sub-optimality in the transfer setting. In this work, we propose a novel algorithm to discover a set of skills, and continuously adapt them along with the higher level even when training on a new task. Our main contributions are two-fold. First, we derive a new hierarchical policy gradient with an unbiased latent-dependent baseline, and we introduce Hierarchical Proximal Policy Optimization (HiPPO), an on-policy method to efficiently train all levels of the hierarchy jointly. Second, we propose a method for training time-abstractions that improves the robustness of the obtained skills to environment changes. Code and results are available at this http URL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fee0d45f-7d64-4081-8535-6d91e234be47Cited by top-tier papers20
- Accelerating Robotic Reinforcement Learning via Parameterized Action PrimitivesMurtaza Dalal, Deepak Pathak, Ruslan SalakhutdinovNeurIPS 2021 · 121 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- Generalized Hindsight for Reinforcement LearningAlexander C. Li, Lerrel Pinto, Pieter AbbeelNeurIPS 2020 · 81 citations
- Hierarchical Skills for Efficient ExplorationJonas Gehring, Gabriel Synnaeve, Andreas Krause, Nicolas UsunierNeurIPS 2021 · 52 citations
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe et al.ICML 2021 · 48 citations
Related papers
- Meta-learning Parameterized SkillsHaotian Fu, Shangqun Yu, Saket Tiwari, Michael Littman et al.ICML 2023 · 8 citations
- DHRL: A Graph-Based Approach for Long-Horizon and Sparse Hierarchical Reinforcement LearningSeungjae Lee, Jigang Kim, Inkyu Jang, H. Jin KimNeurIPS 2022 · 33 citations
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement LearningJinmin He, Kai Li, Yifan Zang, Haobo Fu et al.ICML 2025
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon ReasoningDhruv Shah, Peng Xu, Yao Lu, Ted Xiao et al.ICLR 2022 · 50 citations
- Toward Robust Long Range Policy TransferWei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, Min SunAAAI 2021 · 8 citations
