Parallelizing Model-based Reinforcement Learning Over the Sequence Length
Zirui Wang, Yue Deng, Junfeng Long, Yin Zhang
Abstract
Recently, Model-based Reinforcement Learning (MBRL) methods have demonstrated stunning sample efficiency in various RL domains. However, achieving this extraordinary sample efficiency comes with additional training costs in terms of computations, memory, and training time. To address these challenges, we propose the Pa rallelized Mo del-based R einforcement L earning ( PaMoRL ) framework. PaMoRL introduces two novel techniques: the P arallel W orld M odel ( PWM ) and the P arallelized E ligibility T race E stimation ( PETE ) to parallelize both model learning and policy learning stages of current MBRL methods over the sequence length. Our PaMoRL framework is hardware-efficient and stable, and it can be applied to various tasks with discrete or continuous action spaces using a single set of hyperparameters. The empirical results demonstrate that the PWM and PETE within PaMoRL significantly increase training speed without sacrificing inference efficiency. In terms of sample efficiency, PaMoRL maintains an MBRL-level sample efficiency that outperforms other no-look-ahead MBRL methods and model-free RL methods, and it even exceeds the performance of planning-based MBRL methods and methods with larger networks in certain tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efb43944-04b7-4034-b1bb-16752a8c584cCited by top-tier papers4
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement LearningWenchang Duan, Yaoliang Yu, Jiwan He, Yi ShiNeurIPS 2025 · 11 citations
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng et al.ICLR 2026 · 5 citations
- EDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence ModelingJia-Hua Lee, Bor-Jiun Lin, Wei-Fang Sun, Chun-Yi LeeNeurIPS 2025 · 4 citations
- Sable: a Performant, Efficient and Scalable Sequence Model for MARLOmayma Mahjoub, Sasha Abramowitz, Ruan John de Kock, Wiem Khlifi et al.ICML 2025
Builds on27
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
Related papers
- TaskLoom: Weaving Knowledge Across Tasks in World ModelsQingzhang Zeng, Peixi Peng, hang li, Luntong Li et al.ICML 2026
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter EfficientWenlong Wang, Ivana Dusparic, Yucheng Shi, Ke Zhang et al.ICLR 2025
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski et al.ICML 2020 · 52 citations
