Bayesian Nonparametrics for Offline Skill Discovery
Valentin Villecroze, Harry J. Braviner, Panteha Naderian, Chris J. Maddison, Gabriel Loaiza-Ganem
Abstract
Skills or low-level policies in reinforcement learning are temporally extended actions that can speed up learning and enable complex behaviours. Recent work in offline reinforcement learning and imitation learning has proposed several techniques for skill discovery from a set of expert trajectories. While these methods are promising, the number K of skills to discover is always a fixed hyperparameter, which requires either prior knowledge about the environment or an additional parameter search to tune it. We first propose a method for offline learning of options (a particular skill framework) exploiting advances in variational inference and continuous relaxations. We then highlight an unexplored connection between Bayesian nonparametrics and offline skill discovery, and show how to obtain a nonparametric version of our model. This version is tractable thanks to a carefully structured approximate posterior with a dynamicallychanging number of options, removing the need to specify K. We also show how our nonparametric extension can be applied in other skill frameworks, and empirically demonstrate that our method can outperform state-of-the-art offline skill learning algorithms across a variety of environments. Our code is available at https: //github.com/layer6ai-labs/BNPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e40a422-2693-4833-bc4c-6d1be10ab02cCited by top-tier papers9
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 80 citations
- Efficient Planning with Latent DiffusionWenhao LiICLR 2024 · 15 citations
- Stylized Offline Reinforcement Learning: Extracting Diverse High-Quality Behaviors from Heterogeneous DatasetsYihuan Mao, Chengjie Wu, Xi Chen, Hao Hu et al.ICLR 2024 · 9 citations
- Verifying the Union of Manifolds Hypothesis for Image DataBradley C. A. Brown, Anthony L. Caterini, Brendan Leigh Ross, Jesse C. Cresswell et al.ICLR 2023 · 6 citations
- Beyond Rewards: a Hierarchical Perspective on Offline Multiagent Behavioral AnalysisShayegan Omidshafiei, Andrei Kapishnikov, Yannick Assogba, Lucas Dixon et al.NeurIPS 2022 · 5 citations
Builds on7
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Learning Robot Skills with Temporal Variational InferenceTanmay Shankar, Abhinav GuptaICML 2020 · 80 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon et al.AAAI 2020 · 51 citations
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient EstimatorMax B. Paulus, Chris J. Maddison, Andreas KrauseICLR 2021 · 48 citations
Related papers
- Meta-learning Parameterized SkillsHaotian Fu, Shangqun Yu, Saket Tiwari, Michael Littman et al.ICML 2023 · 8 citations
- Unsupervised Skill Discovery with Bottleneck Option LearningJaekyeom Kim, Seohong Park, Gunhee KimICML 2021 · 39 citations
- Reasoning with Latent Diffusion in Offline Reinforcement LearningSiddarth Venkatraman, Shivesh Khaitan, Ravi Tej Akella, John M. Dolan et al.ICLR 2024 · 39 citations
- Learning Options via CompressionYiding Jiang, Evan Zheran Liu, Benjamin Eysenbach, J. Zico Kolter et al.NeurIPS 2022 · 26 citations
- Data-efficient Hindsight Off-policy Option LearningMarkus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe et al.ICML 2021 · 48 citations
