DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving
Chen Shi, Shaoshuai Shi, Kehua Sheng, Bo Zhang, Li Jiang
Abstract
Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present DriveX, a self-supervised world model that learns generalizable scene dynamics and holistic representations (geometric, semantic, and motion) from large-scale driving videos. DriveX introduces Omni Scene Modeling (OSM), a module that unifies multimodal supervision-3D point cloud forecasting, 2D semantic representation, and image generation-to capture comprehensive scene evolution. To simplify learning complex dynamics, we propose a decoupled latent world modeling strategy that separates world representation learning from future state decoding, augmented by dynamic-aware ray sampling to enhance motion modeling. For downstream adaptation, we design Future Spatial Attention (FSA), a unified paradigm that dynamically aggregates spatiotemporal features from DriveX's predictions to enhance task-specific inference. Extensive experiments demonstrate DriveX's effectiveness: it achieves significant improvements in 3D future point cloud prediction over prior work, while attaining state-of-the-art results on diverse tasks including occupancy prediction, flow estimation, and end-to-end driving. These results validate DriveX's capability as a general-purpose world model, paving the way for robust and unified autonomous driving frameworks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d324a701-32b4-4a72-8e59-787fd2d05543Cited by top-tier papers4
- Driving on RegistersEllington Kirby, Alexandre Boulch, Yihong Xu, Yuan Yin et al.CVPR 2026 · 45 citations
- Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-TuningMuleilan Pei, Shaoshuai Shi, Shaojie ShenICLR 2026 · 21 citations
- WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous DrivingPengxuan Yang, Ben Lu, Zhongpu Xia, Chao Han et al.AAAI 2026 · 8 citations
- GEM: Generating LiDAR World Model via Deformable MambaYang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu et al.CVPR 2026 · 1 citation
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel et al.ICCV 2023 · 800 citations
Related papers
- DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous DrivingYiyao Zhu, Ying Xue, Haiming Zhang, Guangfeng Jiang et al.CVPR 2026 · 2 citations
- DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingChen Min, Dawei Zhao, Liang Xiao, Jian Zhao et al.CVPR 2024 · 20 citations
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du et al.AAAI 2025 · 53 citations
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationXin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen et al.ICCV 2025 · 5 citations
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene ReconstructionChen Shi, Shaoshuai Shi, Xiaoyang Lyu, Chunyang Liu et al.ICLR 2026 · 10 citations
