World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World Model
Yupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang, Yuhang Zheng, Yinfeng Gao, Pengfei Li, Teng Zhang, Zhongpu Xia, Peng Jia, Xianpeng Lang, Dongbin Zhao
Abstract
End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: constructing an informative driving world model to enable perception annotation-free, end-to-end planning via self-supervised learning. In this paper, we present World4Drive, an end-to-end autonomous driving framework that employs vision foundation models to build latent world models for generating and evaluating multi-modal planning trajectories. Specifically, World4Drive first extracts scene features, including driving intention and world latent representations enriched with spatial-semantic priors provided by vision foundation models. It then generates multi-modal planning trajectories based on current scene features and driving intentions and predicts multiple intention-driven future states within the latent space. Finally, it introduces a world model selector module to evaluate and select the best trajectory. We achieve perception annotation-free, end-to-end planning through self-supervised alignment between actual future observations and predicted observations reconstructed from the latent space. World4Drive achieves state-of-the-art performance without manual perception annotations on both the open-loop nuScenes and closed-loop NavSim benchmarks, demonstrating an 18.1% relative reduction in L2 error, 46.7% lower collision rate, and 3.75 faster training convergence. Codes will be accessed at https://github.com/ucaszyp/World4Drive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b36db4b0-e9cb-4dbf-a3ff-6191365bb17cCited by top-tier papers14
- Driving on RegistersEllington Kirby, Alexandre Boulch, Yihong Xu, Yuan Yin et al.CVPR 2026 · 45 citations
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous DrivingFeiyang Jia, Lin Liu, Ziying Song, Caiyan Jia et al.ICML 2026 · 20 citations
- MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous DrivingJunli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing et al.CVPR 2026 · 16 citations
- ResWorld: Temporal Residual World Model for End-to-End Autonomous DrivingJinqing Zhang, Zehua Fu, Zelin Xu, Wenying Dai et al.ICLR 2026 · 13 citations
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationZhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou et al.CVPR 2026 · 12 citations
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai et al.ICCV 2023 · 388 citations
Related papers
- Enhancing End-to-End Autonomous Driving with Latent World ModelYingyan Li, Lue Fan, Jiawei He, Yuqi Wang et al.ICLR 2025
- WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous DrivingPengxuan Yang, Ben Lu, Zhongpu Xia, Chao Han et al.AAAI 2026 · 8 citations
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du et al.AAAI 2025 · 53 citations
- End-to-End Driving with Online Trajectory Evaluation via BEV World ModelYingyan Li, Yuqi Wang, Yang Liu, Jiawei He et al.ICCV 2025 · 17 citations
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationYichen Xie, Runsheng Xu, Tong He, Jyh-Jing Hwang et al.CVPR 2025
