DriveLaW: Unifying Planning and Video Generation in a Latent Driving World
Tianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao, Kaixin Xiong, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang
Abstract
World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate within ostensibly unified architectures that still keep world prediction and motion planning as decoupled processes. To bridge this gap, we propose DriveLaW, a novel paradigm that unifies video generation and motion planning. By directly injecting the latent representation from its video generator into the planner, DriveLaW ensures inherent consistency between high-fidelity future generation and reliable trajectory planning. Specifically, DriveLaW consists of two core components: DriveLaW-Video, our powerful world model that generates high-fidelity forecasting with expressive latent representations, and DriveLaW-Act, a diffusion planner that generates consistent and reliable trajectories from the latent of DriveLaW-Video, with both components optimized by a three-stage progressive training strategy. New state-of-the-art results across both tasks demonstrate the power of our unified paradigm. DriveLaW not only significantly advances video prediction, surpassing the previous best-performing work by 33.3% in FID and 1.8% in FVD, but also sets a new record on the NAVSIM planning benchmark. Code available at https://github.com/xiaomi- research/drivelaw.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8720059d-5549-4bfc-be2c-4becc43cd3d4Builds on42
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
Related papers
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous DrivingFeiyang Jia, Lin Liu, Ziying Song, Caiyan Jia et al.ICML 2026 · 20 citations
- From Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionZhida Zhao, Talas Fu, Yifan Wang, Lijun Wang et al.NeurIPS 2025 · 36 citations
- Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World ModelXiaodong Wang, Zhirong Wu, Peixi PengAAAI 2026
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationGuosheng Zhao, Xiaofeng Wang, Zheng Zhu, Xinze Chen et al.AAAI 2025 · 31 citations
- Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent SpaceJian Zhu, Zhengyu Jia, Tian Gao, Jiaxin Deng et al.AAAI 2026 · 5 citations
