Unsupervised Hierarchical Skill Discovery
Damion Harvey, Geraud Nangue Tasse, Benjamin Rosman, Branden Ingram, Steven James
Abstract
We consider the problem of unsupervised skill segmentation and hierarchical structure discovery in reinforcement learning. While recent approaches have sought to segment trajectories into reusable skills or options, most rely on action labels, rewards, or handcrafted annotations, limiting their applicability. We propose a method that segments unlabelled trajectories into skills and induces a hierarchical structure over them using a grammar-based approach. The resulting hierarchy captures both low-level behaviours and their composition into higher-level skills. We evaluate our approach in high-dimensional, pixel-based environments, including Craftax and the full, unmodified version of Minecraft. Using metrics for skill segmentation, reuse, and hierarchy quality, we find that our method consistently produces more structured and semantically meaningful hierarchies than existing baselines. Furthermore, as a proof of concept, we demonstrate that these discovered hierarchies accelerate and stabilise learning on downstream reinforcement learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bfd989c-0c8e-48a0-a782-ad4f3b8b29d8Builds on4
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement LearningMichael T. Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan et al.ICML 2024 · 71 citations
- Learning Task Decomposition with Ordered Memory Policy NetworkYuchen Lu, Yikang Shen, Siyuan Zhou, Aaron C. Courville et al.ICLR 2021 · 17 citations
- Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action SegmentationMing Xu, Stephen GouldCVPR 2024 · 15 citations
Related papers
- Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement LearningJiaheng Hu, Zizhao Wang, Peter Stone, Roberto Martín-MartínNeurIPS 2024 · 21 citations
- Hierarchical Multi-Agent Skill DiscoveryMingyu Yang, Yaodong Yang, Zhenbo Lu, Wengang Zhou et al.NeurIPS 2023 · 34 citations
- Language-guided Skill Learning with Temporal Variational InferenceHaotian Fu, Pratyusha Sharma, Elias Stengel-Eskin, George Konidaris et al.ICML 2024 · 11 citations
- Creating Multi-Level Skill Hierarchies in Reinforcement LearningJoshua B. Evans, Özgür SimsekNeurIPS 2023 · 15 citations
- Skill Discovery for Exploration and Planning using Deep Skill GraphsAkhil Bagaria, Jason K. Senthil, George KonidarisICML 2021 · 73 citations
