DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving
Zhenhua Xu, Yan Bai, Yujia Zhang, Zhuoling Li, Fei Xia, Kwan-Yee K. Wong, Jianqiang Wang, Hengshuang Zhao
Abstract
Multimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which focused on open-loop tasks, this study explores the capabilities of LLMs in enhancing closed-loop autonomous driving. DriveGPT4-V2 processes camera images and vehicle states as input to generate lowlevel control signals for end-to-end vehicle operation. A multi-view visual tokenizer (MV-VT) is employed enabling DriveGPT4-V2 to perceive the environment with an extensive range while maintaining critical details. The model architecture has been refined to improve decision prediction and inference speed. To further enhance the performance, an additional expert LLM is trained for online imitation learning. The expert LLM, sharing a similar structure with DriveGPT4-V2, can access privileged information about surrounding objects for more robust and reliable predictions. Experimental results show that DriveGPT4-V2 outperforms all baselines on the challenging CARLA Longest6 benchmark. The code and data of DriveGPT4-V2 will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 119a181c-347b-4027-a0e7-3ce36ad977c4Cited by top-tier papers5
- BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous DrivingShu Liu, Wenlin Chen, Weihao Li, Zheng Wang et al.ICLR 2026 · 19 citations
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and PlanningZhe Liu, Runhui Huang, Rui Yang, Siming Yan et al.CVPR 2026 · 15 citations
- CausalVAD: De-confounding End-to-End Autonomous Driving via Causal InterventionJiacheng Tang, Zhiyuan Zhou, Zhuolin He, Jia Zhang et al.CVPR 2026 · 8 citations
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language ModelsYuqi Liu, Liangyu Chen, Jiazhen Liu, Mingkang Zhu et al.ICML 2026
- LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly DetectionZhichao Zeng, Jiasheng Zhang, Jiyun Sun, Jiangtao Cui et al.CVPR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
Related papers
- LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsHao Shao, Yuxuan Hu, Letian Wang, Guanglu Song et al.CVPR 2024 · 114 citations
- Driving with Advice: Large Model as Motion Advisor for Joint PlanningJunyin Wang, Jinlei Yu, Hao Lin, Huikai Liu et al.AAAI 2026
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- Orion: A Holistic End-To-End Autonomous Driving Framework by Vision-Language Instructed Action GenerationHaoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui et al.ICCV 2025 · 19 citations
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar et al.ICLR 2024 · 143 citations
