COME: Adding Scene-Centric Forecasting Control to Occupancy World Model
Yining Shi, Kun Jiang, Qiang Meng, Ke Wang, Jiabao Wang, Wenchao Sun, Tuopu Wen, Mengmeng Yang, Diange Yang
Abstract
World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to separate environmental changes from ego-motion by leveraging the scene-centric coordinate systems. In this paper, we introduce COME: a framework that integrates scene-centric forecasting Control into the Occupancy world ModEl. Specifically, COME first generates ego-irrelevant, spatially consistent future features through a scene-centric prediction branch, which are then converted into scene condition using a tailored ControlNet. These condition features are subsequently injected into the occupancy world model, enabling more accurate and controllable future occupancy predictions. Experimental results on the nuScenes-Occ3D dataset show that COME achieves consistent and significant improvements over state-of-the-art (SOTA) methods across diverse configurations, including different input sources (ground-truth, camera-based, fusion-based occupancy) and prediction horizons (3s and 8s). For example, under the same settings, COME achieves 26.3% better mIoU metric than DOME and 23.7% better mIoU metric than UniScene. These results highlight the efficacy of disentangled representation learning in enhancing spatio-temporal prediction fidelity for world models. Code and videos will be available at https://github.com/synsin0/COME.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationZhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou et al.CVPR 2026 · 12 citations
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World ModelJiayuan Du, Yiming Zhao, Zhenglong Guo, Yong Pan et al.CVPR 2026 · 6 citations
- GEM: Generating LiDAR World Model via Deformable MambaYang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu et al.CVPR 2026 · 1 citation
- Grounded Latents for Entity-Centric 4D Scene GenerationJinhyung Park, Navyata Sanghvi, Erica Weng, Shawn Hunt et al.CVPR 2026
Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du et al.AAAI 2025 · 53 citations
- Visual Point Cloud Forecasting Enables Scalable Autonomous DrivingZetong Yang, Li Chen, Yanan Sun, Hongyang LiCVPR 2024 · 40 citations
Related papers
- Uniocc: a Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous DrivingYuping Wang, Xiangyu Huang, Xiaokang Sun, Mingxuan Yan et al.ICCV 2025 · 3 citations
- Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous DrivingXiang Li, Pengfei Li, Yupeng Zheng, Wei Sun et al.ICLR 2025
- Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsJunyi Ma, Xieyuanli Chen, Jiawei Huang, Jingyi Xu et al.CVPR 2024 · 28 citations
- DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingChen Min, Dawei Zhao, Liang Xiao, Jian Zhao et al.CVPR 2024 · 20 citations
- CSV-Occ: Fusing Multi-frame Alignment for Occupancy Prediction with Temporal Cross State Space Model and Central Voting MechanismZiming Zhu, Yu Zhu, Jiahao Chen, Xiaofeng Ling et al.ICML 2025
