Sparse Imagination for Efficient Visual World Model Planning
Junha Chun, Youngjoon Jeong, Taesup Kim
Abstract
World model based planning has significantly improved decision-making in complex environments by enabling agents to simulate future states and make informed choices. This computational burden is particularly restrictive in robotics, where resources are severely constrained. To address this limitation, we propose a Sparse Imagination for Efficient Visual World Model Planning, which enhances computational efficiency by reducing the number of tokens processed during forward prediction. Our method leverages a sparsely trained vision-based world model based on transformers with randomized grouped attention strategy, allowing the model to flexibly adjust the number of tokens processed based on the computational resource. By enabling sparse imagination during latent rollout, our approach significantly accelerates planning while maintaining high control fidelity. Experimental results demonstrate that sparse imagination preserves task performance while dramatically improving inference efficiency. This general technique for visual planning is applicable from simple test-time trajectory optimization to complex real-world tasks with the latest VLAs, enabling the deployment of world models in real-time scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eca02cd0-f2e2-4cc0-ae5a-7a164bdcbd1bCited by top-tier papers2
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu et al.NeurIPS 2025 · 12 citations
- DDP-WM: Disentangled Dynamics Prediction for Efficient World ModelsShicheng Yin, Kaixuan Yin, Weixing Chen, Yang Liu et al.ICML 2026 · 3 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World ModelDongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho et al.CVPR 2026 · 6 citations
- MaskViT: Masked Visual Pre-Training for Video PredictionAgrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu et al.ICLR 2023 · 45 citations
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World ModelJiayuan Du, Yiming Zhao, Zhenglong Guo, Yong Pan et al.CVPR 2026 · 6 citations
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 44 citations
- DreamPhase: Offline Imagination and Uncertainty-Guided Planning for Large-Language-Model AgentsShayan Mohajer Hamidi, Linfeng Ye, Konstantinos N. PlataniotisICLR 2026
