Effectively Learning Initiation Sets in Hierarchical Reinforcement Learning
Akhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matt Corsaro, Sreehari Rammohan, George Dimitri Konidaris
摘要
An agent learning an option in hierarchical reinforcement learning must solve three problems: identify the option's subgoal (termination condition), learn a policy, and learn where that policy will succeed (initiation set). The termination condition is typically identified first, but the option policy and initiation set must be learned simultaneously, which is challenging because the initiation set depends on the option policy, which changes as the agent learns. Consequently, data obtained from option execution becomes invalid over time, leading to an inaccurate initiation set that subsequently harms downstream task performance. We highlight three issues-data non-stationarity, temporal credit assignment, and pessimism-specific to learning initiation sets, and propose to address them using tools from off-policy value estimation and classification. We show that our method learns higher-quality initiation sets faster than existing methods (in MINIGRID and MONTEZUMA'S REVENGE), can automatically discover promising grasps for robot manipulation (in ROBOSUITE), and improves the performance of a state-of-the-art option discovery method in a challenging maze navigation task in MuJoCo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 被引用 114 次
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 被引用 19 次
- SkiLD: Unsupervised Skill Discovery Guided by Factor InteractionsZizhao Wang, Jiaheng Hu, Caleb Chuck, Stephen Chen 等NeurIPS 2024 · 被引用 15 次
- Leveraging Skills from Unlabeled Prior Data for Efficient Online ExplorationMax Wilcoxson, Qiyang Li, Kevin Frans, Sergey LevineICML 2025
- Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement LearningYiming Fei, Ziming Wang, Rui Yan, Huajin TangICML 2026
它引用的顶会 Paper4
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 被引用 126 次
- Skill Discovery for Exploration and Planning using Deep Skill GraphsAkhil Bagaria, Jason K. Senthil, George KonidarisICML 2021 · 被引用 73 次
- What can I do here? A Theory of Affordances in Reinforcement LearningKhimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, David Abel 等ICML 2020 · 被引用 60 次
- On Efficiency in Hierarchical Reinforcement LearningZheng Wen, Doina Precup, Morteza Ibrahimi, André Barreto 等NeurIPS 2020 · 被引用 44 次
相关 Paper
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 被引用 38 次
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu 等NeurIPS 2021 · 被引用 38 次
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe 等ICML 2021 · 被引用 48 次
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon 等AAAI 2020 · 被引用 51 次
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 被引用 97 次
