SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
Jiayuan Du, Yiming Zhao, Zhenglong Guo, Yong Pan, Wenbo Hou, Zhihui Hao, Kun Zhan, Qijun Chen
Abstract
This paper introduces a novel architecture for trajectoryconditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs) to generate discrete occupancy tokens, which inherently limit representational capacity, our approach predicts multi-frame future occupancy in an end-to-end manner directly from raw image features. Inspired by the success of attention-based transformer architectures in foundational vision and language models such as GPT and VGGT, we employ a sparse occupancy representation that bypasses the intermediate bird's eye view (BEV) projection and its explicit geometric priors. This design allows the transformer to capture spatiotemporal dependencies more effectively. By avoiding both the finite-capacity constraint of discrete tokenization and the structural limitations of BEV representations, our method achieves state-of-the-art performance on the nuScenes benchmark for 1-3 second occupancy forecasting, outperforming existing approaches by a significant margin. Furthermore, it demonstrates robust scene dynamics understanding, consistently delivering high accuracy under arbitrary future trajectory conditioning. Code: https://github.com/MrPicklesGG/SparseWorld.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68602bfb-f8a0-4a59-af4b-dd7bf1960ed4Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
Related papers
- GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial UnderstandingHaoyi Jiang, Liu Liu, Tianheng Cheng, Xinjie Wang et al.CVPR 2025
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou et al.CVPR 2023
- Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy PredictionChi Yan, Dan XuICLR 2026 · 6 citations
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionPin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng et al.CVPR 2024 · 37 citations
