Data-efficient Hindsight Off-policy Option Learning
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Y. Siegel, Nicolas Heess, Martin A. Riedmiller
摘要
Solutions to most complex tasks can be decomposed into simpler, intermediate skills, reusable across wider ranges of problems. We follow this concept and introduce Hindsight Off-policy Options (HO2), a new algorithm for efficient and robust option learning. The algorithm relies on critic-weighted maximum likelihood estimation and an efficient dynamic programming inference procedure over off-policy trajectories. We can backpropagate through the inference procedure through time and the policy components for every time-step, making it possible to train all component's parameters off-policy, independently of the data-generating behavior policy. Experimentally, we demonstrate that HO2 outperforms competitive baselines and solves demanding robot stacking and ball-in-cup tasks from raw pixel inputs in simulation. We further compare autoregressive option policies with simple mixture policies, providing insights into the relative impact of two types of abstractions common in the options framework: action abstraction and temporal abstraction. Finally, we illustrate challenges caused by stale data in off-policy options learning and provide effective solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 被引用 173 次
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 被引用 80 次
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 被引用 38 次
- Learning transferable motor skills with hierarchical latent mixture policiesDushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier 等ICLR 2022 · 被引用 34 次
- Solving Compositional Reinforcement Learning Problems via Task ReductionYunfei Li, Yilin Wu, Huazhe Xu, Xiaolong Wang 等ICLR 2021 · 被引用 20 次
它引用的顶会 Paper1
相关 Paper
- Learning Robot Skills with Temporal Variational InferenceTanmay Shankar, Abhinav GuptaICML 2020 · 被引用 80 次
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 被引用 97 次
- Adversarial Option-Aware Hierarchical Imitation LearningMingxuan Jing, Wenbing Huang, Fuchun Sun, Xiaojian Ma 等ICML 2021 · 被引用 28 次
- Bayesian Nonparametrics for Offline Skill DiscoveryValentin Villecroze, Harry J. Braviner, Panteha Naderian, Chris J. Maddison 等ICML 2022 · 被引用 9 次
- Learning Temporally AbstractWorld Models without Online ExperimentationBenjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider 等ICML 2023 · 被引用 7 次
