World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World Model
Yupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang, Yuhang Zheng, Yinfeng Gao, Pengfei Li, Teng Zhang, Zhongpu Xia, Peng Jia, Xianpeng Lang, Dongbin Zhao
摘要
End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: constructing an informative driving world model to enable perception annotation-free, end-to-end planning via self-supervised learning. In this paper, we present World4Drive, an end-to-end autonomous driving framework that employs vision foundation models to build latent world models for generating and evaluating multi-modal planning trajectories. Specifically, World4Drive first extracts scene features, including driving intention and world latent representations enriched with spatial-semantic priors provided by vision foundation models. It then generates multi-modal planning trajectories based on current scene features and driving intentions and predicts multiple intention-driven future states within the latent space. Finally, it introduces a world model selector module to evaluate and select the best trajectory. We achieve perception annotation-free, end-to-end planning through self-supervised alignment between actual future observations and predicted observations reconstructed from the latent space. World4Drive achieves state-of-the-art performance without manual perception annotations on both the open-loop nuScenes and closed-loop NavSim benchmarks, demonstrating an 18.1% relative reduction in L2 error, 46.7% lower collision rate, and 3.75 faster training convergence. Codes will be accessed at https://github.com/ucaszyp/World4Drive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Driving on RegistersEllington Kirby, Alexandre Boulch, Yihong Xu, Yuan Yin 等CVPR 2026 · 被引用 45 次
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous DrivingFeiyang Jia, Lin Liu, Ziying Song, Caiyan Jia 等ICML 2026 · 被引用 20 次
- MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous DrivingJunli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing 等CVPR 2026 · 被引用 16 次
- ResWorld: Temporal Residual World Model for End-to-End Autonomous DrivingJinqing Zhang, Zehua Fu, Zelin Xu, Wenying Dai 等ICLR 2026 · 被引用 13 次
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationZhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta 等NeurIPS 2024 · 被引用 403 次
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 等ICCV 2023 · 被引用 388 次
相关 Paper
- Enhancing End-to-End Autonomous Driving with Latent World ModelYingyan Li, Lue Fan, Jiawei He, Yuqi Wang 等ICLR 2025
- WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous DrivingPengxuan Yang, Ben Lu, Zhongpu Xia, Chao Han 等AAAI 2026 · 被引用 8 次
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingYu Yang, Jianbiao Mei, Yukai Ma, Siliang Du 等AAAI 2025 · 被引用 53 次
- End-to-End Driving with Online Trajectory Evaluation via BEV World ModelYingyan Li, Yuqi Wang, Yang Liu, Jiawei He 等ICCV 2025 · 被引用 17 次
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationYichen Xie, Runsheng Xu, Tong He, Jyh-Jing Hwang 等CVPR 2025
