Data-efficient Hindsight Off-policy Option Learning
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Y. Siegel, Nicolas Heess, Martin A. Riedmiller
Abstract
Solutions to most complex tasks can be decomposed into simpler, intermediate skills, reusable across wider ranges of problems. We follow this concept and introduce Hindsight Off-policy Options (HO2), a new algorithm for efficient and robust option learning. The algorithm relies on critic-weighted maximum likelihood estimation and an efficient dynamic programming inference procedure over off-policy trajectories. We can backpropagate through the inference procedure through time and the policy components for every time-step, making it possible to train all component's parameters off-policy, independently of the data-generating behavior policy. Experimentally, we demonstrate that HO2 outperforms competitive baselines and solves demanding robot stacking and ball-in-cup tasks from raw pixel inputs in simulation. We further compare autoregressive option policies with simple mixture policies, providing insights into the relative impact of two types of abstractions common in the options framework: action abstraction and temporal abstraction. Finally, we illustrate challenges caused by stale data in off-policy options learning and provide effective solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c96489e8-bed9-4b07-a144-032e53f8d0edCited by top-tier papers17
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 80 citations
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 38 citations
- Learning transferable motor skills with hierarchical latent mixture policiesDushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier et al.ICLR 2022 · 34 citations
- Solving Compositional Reinforcement Learning Problems via Task ReductionYunfei Li, Yilin Wu, Huazhe Xu, Xiaolong Wang et al.ICLR 2021 · 20 citations
Builds on1
Related papers
- Learning Robot Skills with Temporal Variational InferenceTanmay Shankar, Abhinav GuptaICML 2020 · 80 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- Adversarial Option-Aware Hierarchical Imitation LearningMingxuan Jing, Wenbing Huang, Fuchun Sun, Xiaojian Ma et al.ICML 2021 · 28 citations
- Bayesian Nonparametrics for Offline Skill DiscoveryValentin Villecroze, Harry J. Braviner, Panteha Naderian, Chris J. Maddison et al.ICML 2022 · 9 citations
- Learning Temporally AbstractWorld Models without Online ExperimentationBenjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider et al.ICML 2023 · 7 citations
