DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive Transformers
Yuntao Chen, Yuqi Wang, Zhaoxiang Zhang
Abstract
World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ec1e3e5-0619-4c83-a584-0de613b9b70cCited by top-tier papers23
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingShuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu et al.NeurIPS 2025 · 228 citations
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous DrivingYongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li et al.ICLR 2026 · 196 citations
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous DrivingYingyan Li, Shuyao Shang, Weisong Liu, Bing Zhan et al.ICLR 2026 · 134 citations
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li et al.ICLR 2026 · 93 citations
- DriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldTianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao et al.CVPR 2026 · 58 citations
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- iVideoGPT: Interactive VideoGPTs are Scalable World ModelsJialong Wu, Shaofeng Yin, Ningya Feng, Xu He et al.NeurIPS 2024 · 177 citations
- DriveGPT: Scaling Autoregressive Behavior Models for DrivingXin Huang, Eric M. Wolff, Paul Vernaza, Tung Phan-Minh et al.ICML 2025
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng et al.ICCV 2023 · 186 citations
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- Diffusion-Based Planning for Autonomous Driving with Flexible GuidanceYinan Zheng, Ruiming Liang, Kexin Zheng, Jinliang Zheng et al.ICLR 2025
