Enhancing Exploration and Exploitation in Hierarchical Reinforcement Learning with Subgoal Graph Learning
Yibo Zhang, Dengpeng Xing
Abstract
Goal-conditioned hierarchical reinforcement learning has demonstrated effectiveness in addressing complicated decision-making tasks by providing "temporal extraction", which decomposes tasks into smaller and more manageable "subgoals". This enables agents to plan over a longer time scale. However, achieving optimal exploration and exploitation still remains a challenge, especially for long-horizon or sparse-reward scenarios. In this paper, we introduce Active exploration and hierarchical Self-Imitation (ASI), an effective scheme to enhance exploration and exploitation based on subgoal representation learning. The key point of ASI is to utilize temporal adjacency information in the representation space. We construct and dynamically update an adjacency graph that captures the relationships between subgoals. Based on the adjacency information provided by the graph, we design two mechanisms: active "frontier-reaching" exploration that faster expands the explored area by targeting boundary regions, and hierarchical self-imitation learning that leverages historical experience to facilitate both frontier reaching and policy training. Experimental results show that our method accelerates exploration and outperforms existing baselines in challenging long-horizon continuous control tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d622b8b6-ae41-4ab3-af84-298ff32c7863Builds on11
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 183 citations
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 152 citations
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie et al.ICML 2020 · 145 citations
- Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement LearningTianren Zhang, Shangqi Guo, Tian Tan, Xiaolin Hu et al.NeurIPS 2020 · 112 citations
- World Model as a Graph: Learning Latent Landmarks for PlanningLunjun Zhang, Ge Yang, Bradly C. StadieICML 2021 · 90 citations
Related papers
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu et al.ICLR 2022 · 28 citations
- Learning Subgoal Representations with Slow DynamicsSiyuan Li, Lulu Zheng, Jianhao Wang, Chongjie ZhangICLR 2021 · 48 citations
- Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement LearningSeungyul Han, Jaebak Hwang, Sanghyeon Lee, Jeongmo KimICLR 2026 · 3 citations
- Imitating Graph-Based Planning with Goal-Conditioned PoliciesJunsu Kim, Younggyo Seo, Sungsoo Ahn, Kyunghwan Son et al.ICLR 2023 · 2 citations
- DHRL: A Graph-Based Approach for Long-Horizon and Sparse Hierarchical Reinforcement LearningSeungjae Lee, Jigang Kim, Inkyu Jang, H. Jin KimNeurIPS 2022 · 33 citations
