Learning Temporally AbstractWorld Models without Online Experimentation
Benjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider, Howie Choset
Abstract
Agents that can build temporally abstract representations of their environment are better able to understand their world and make plans on extended time scales, with limited computational power and modeling capacity. However, existing methods for automatically learning temporally abstract world models usually require millions of online environmental interactions and incentivize agents to reach every accessible environmental state, which is infeasible for most realworld robots both in terms of data efficiency and hardware safety. In this paper, we present an approach for simultaneously learning sets of skills and temporally abstract, skill-conditioned world models purely from offline data, enabling agents to perform zero-shot online planning of skill sequences for new tasks. We show that our approach performs comparably to or better than a wide array of state-of-the-art offline RL algorithms on a number of simulated robotics locomotion and manipulation benchmarks, while offering a higher degree of adaptability to new goals. Finally, we show that our approach offers a much higher degree of robustness to perturbations in environmental dynamics, compared to policy-based methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Latent Plan Transformer for Trajectory Abstraction: Planning as Latent Space InferenceDeqian Kong, Dehong Xu, Minglu Zhao, Bo Pang et al.NeurIPS 2024 · 15 citations
- Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RLJinwoo Choi, Sang-Hyun Lee, Seung-Woo SeoICML 2026 · 3 citations
- Dynamic Contrastive Skill Learning with State-Transition Based Skill Clustering and Dynamic Length AdjustmentJinwoo Choi, Seung-Woo SeoICLR 2025
Builds on8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
Related papers
- Learning transferable motor skills with hierarchical latent mixture policiesDushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier et al.ICLR 2022 · 34 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic SkillsYevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao et al.ICML 2021 · 173 citations
- Skill Machines: Temporal Logic Skill Composition in Reinforcement LearningGeraud Nangue Tasse, Devon Jarvis, Steven James, Benjamin RosmanICLR 2024 · 12 citations
- Hierarchical Few-Shot Imitation with Skill Transition ModelsKourosh Hakhamaneshi, Ruihan Zhao, Albert Zhan, Pieter Abbeel et al.ICLR 2022 · 51 citations
