-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
Zhimin Liao, Ping Wei, Ruijie Zhang, Shuaijia Chen, Haoxuan Wang, Ziyang Ren
Abstract
Forecasting the evolution of 3D scenes and generating unseen scenarios via occupancy-based world models offer substantial potential for addressing corner cases in autonomous driving systems. While tokenization has revolutionized image and video generation, efficiently tokenizing complex 3D scenes remains a critical challenge for 3D world models. To address this issue, we propose -World, an efficient framework for 4D occupancy forecasting. Our method decouples scene tokenization into intra-scene and inter-scene tokenizers. The intra-scene tokenizer employs a multi-scale residual quantization strategy to hierarchically compress scenes while preserving spatial details. The inter-scene tokenizer residually aggregates temporal dependencies across timesteps. This dual design preserves the compactness of 3D tokenizers while retaining the dynamic expressiveness of 4D tokenizers. Unlike decoder-only GPT-style autoregressive models, -World adopts an encoder-decoder architecture. The encoder aggregates spatial context from the current scene and predicts a transformation matrix to enable high-level control over scene generation. The decoder, conditioned on transformation matrix and historical tokens, ensures temporal consistency during generation. Experiments demonstrate that -World achieves state-of-theart performance, outperforming existing methods by in and 36.9% in for 4D occupancy forecasting while exhibiting exceptional computational efficiency. It nearly requires 2.9 GB of training memory and achieves real-time inference at 37.0 FPS. Our code is available on https://github.com/lzzzzzm/II-World.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89aa15f3-db5c-4683-8c72-4f0285115fc9Cited by top-tier papers2
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World ModelJiayuan Du, Yiming Zhao, Zhenglong Guo, Yong Pan et al.CVPR 2026 · 6 citations
- The Structure-Equivalent Prior: Unifying Temporal Dynamics and 3D Evolution in 4D Latent SpaceJingyuan Gao, Tianyu Shen, Ruosen Hao, Te Guo et al.AAAI 2026
Builds on28
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
Related papers
- SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic QueriesChenxu Dang, Haiyan Liu, Jason Bao, Pei An et al.AAAI 2026 · 6 citations
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du et al.AAAI 2025 · 53 citations
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionLunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas et al.ICLR 2024 · 105 citations
- DIO: Decomposable Implicit 4D Occupancy-Flow World ModelChristopher Diehl, Quinlan Sykora, Ben Agro, Thomas Gilles et al.CVPR 2025
- Overcoming Challenges of Long-Horizon Prediction in Driving World ModelsArian Mousakhan, Sudhanshu Mittal, Silvio Galesso, Karim Farid et al.NeurIPS 2025 · 1 citation
