Bootstrap Off-policy with World Model
Guojian Zhan, Likun Wang, Xiangteng Zhang, Jiaxin Gao, Masayoshi Tomizuka, Shengbo Eben Li
Abstract
Online planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevitably introduces a divergence between the collected data and the policy's actual behaviors, degrading both model learning and policy improvement. To address this, we propose BOOM (Bootstrap Off-policy with WOrld Model), a framework that tightly integrates planning and off-policy learning through a bootstrap loop: the policy initializes the planner, and the planner refines actions to bootstrap the policy through behavior alignment. This loop is supported by a jointly learned world model, which enables the planner to simulate future trajectories and provides value targets to facilitate policy improvement. The core of BOOM is a likelihood-free alignment loss that bootstraps the policy using the planner's non-parametric action distribution, combined with a soft value-weighted mechanism that prioritizes high-return behaviors and mitigates variability in the planner's action quality within the replay buffer. Experiments on the high-dimensional DeepMind Control Suite and Humanoid-Bench show that BOOM achieves state-of-the-art results in both training stability and final performance. The code is accessible at https://github.com/molumitu/BOOM_MBRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Langevin Rollout Optimization for Modelic Reinforcement LearningTianyi Zhang, Likun Wang, Guojian Zhan, Feihong Zhang et al.ICML 2026 · 7 citations
- The Surprising Difficulty of Search in Model-Based Reinforcement LearningWei-Di Chang, Mikael Henaff, Brandon Amos, Gregory Dudek et al.ICML 2026 · 4 citations
- Harmonized Dual Policy Improvement for Modelic Reinforcement LearningGuojian Zhan, Likun Wang, Feihong Zhang, Yang Guan et al.ICML 2026
- A KL-regularization framework for learning to plan with adaptive priorsÁlvaro Serra-Gómez, Daniel Jarne Ornia, Dhruva Tirumala, Thomas M MoerlandICML 2026
- Self-supervised Hierarchical Visual Reasoning with World ModelYuanfei Xu, Lin Liu, Wengang Zhou, Mingxiao Feng et al.ICML 2026
Builds on19
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
Related papers
- Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement LearningJiayu Chen, Le Xu, Aravind Venugopal, Jeff SchneiderICML 2026
- Bootstrapped Model Predictive ControlYuhang Wang, Hanwei Guo, Sizhe Wang, Long Qian et al.ICLR 2025
- TaskLoom: Weaving Knowledge Across Tasks in World ModelsQingzhang Zeng, Peixi Peng, hang li, Luntong Li et al.ICML 2026
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng et al.ICLR 2022 · 173 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
