Hieros: Hierarchical Imagination on Structured State Space Sequence World Models
Paul Mattes, Rainer Schlosser, Ralf Herbrich
Abstract
One of the biggest challenges to modern deep reinforcement learning (DRL) algorithms is sample efficiency. Many approaches learn a world model in order to train an agent entirely in imagination, eliminating the need for direct environment interaction during training. However, these methods often suffer from either a lack of imagination accuracy, exploration capabilities, or runtime efficiency. We propose Hieros, a hierarchical policy that learns time abstracted world representations and imagines trajectories at multiple time scales in latent space. Hieros uses an S5 layer-based world model, which predicts next world states in parallel during training and iteratively during environment interaction. Due to the special properties of S5 layers, our method can train in parallel and predict next world states iteratively during imagination. This allows for more efficient training than RNN-based world models and more efficient imagination than Transformer-based world models. We show that our approach outperforms the state of the art in terms of mean and median normalized human score on the Atari 100k benchmark, and that our proposed world model is able to predict complex dynamics very accurately. We also show that Hieros displays superior exploration capabilities compared to existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 15 citations
- GLAM: Global-Local Variation Awareness in Mamba-based World ModelQian He, Wenqi Liang, Chunhui Hao, Gan Sun et al.AAAI 2025 · 2 citations
- DyMoDreamer: World Modeling with Dynamic ModulationBoxuan Zhang, Runqing Wang, Wei Xiao, Weipu Zhang et al.NeurIPS 2025 · 2 citations
Builds on20
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
Related papers
- Transformers are Sample-Efficient World ModelsVincent Micheli, Eloi Alonso, François FleuretICLR 2023 · 11 citations
- Transformer-based World Models Are Happy With 100k InteractionsJan Robine, Marc Höftmann, Tobias Uelwer, Stefan HarmelingICLR 2023 · 4 citations
- Facing Off World Model Backbones: RNNs, Transformers, and S4Fei Deng, Junyeong Park, Sungjin AhnNeurIPS 2023 · 53 citations
- Deep Hierarchical Planning from PixelsDanijar Hafner, Kuang-Huei Lee, Ian Fischer, Pieter AbbeelNeurIPS 2022 · 153 citations
- Learning to Play Atari in a World of TokensPranav Agarwal, Sheldon Andrews, Samira Ebrahimi KahouICML 2024 · 6 citations
