DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving
Zhenhua Xu, Yan Bai, Yujia Zhang, Zhuoling Li, Fei Xia, Kwan-Yee K. Wong, Jianqiang Wang, Hengshuang Zhao
摘要
Multimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which focused on open-loop tasks, this study explores the capabilities of LLMs in enhancing closed-loop autonomous driving. DriveGPT4-V2 processes camera images and vehicle states as input to generate lowlevel control signals for end-to-end vehicle operation. A multi-view visual tokenizer (MV-VT) is employed enabling DriveGPT4-V2 to perceive the environment with an extensive range while maintaining critical details. The model architecture has been refined to improve decision prediction and inference speed. To further enhance the performance, an additional expert LLM is trained for online imitation learning. The expert LLM, sharing a similar structure with DriveGPT4-V2, can access privileged information about surrounding objects for more robust and reliable predictions. Experimental results show that DriveGPT4-V2 outperforms all baselines on the challenging CARLA Longest6 benchmark. The code and data of DriveGPT4-V2 will be publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous DrivingShu Liu, Wenlin Chen, Weihao Li, Zheng Wang 等ICLR 2026 · 被引用 19 次
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and PlanningZhe Liu, Runhui Huang, Rui Yang, Siming Yan 等CVPR 2026 · 被引用 15 次
- CausalVAD: De-confounding End-to-End Autonomous Driving via Causal InterventionJiacheng Tang, Zhiyuan Zhou, Zhuolin He, Jia Zhang 等CVPR 2026 · 被引用 8 次
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language ModelsYuqi Liu, Liangyu Chen, Jiazhen Liu, Mingkang Zhu 等ICML 2026
- LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly DetectionZhichao Zeng, Jiasheng Zhang, Jiyun Sun, Jiangtao Cui 等CVPR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
相关 Paper
- LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsHao Shao, Yuxuan Hu, Letian Wang, Guanglu Song 等CVPR 2024 · 被引用 114 次
- Driving with Advice: Large Model as Motion Advisor for Joint PlanningJunyin Wang, Jinlei Yu, Hao Lin, Huikai Liu 等AAAI 2026
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma 等AAAI 2026 · 被引用 119 次
- Orion: A Holistic End-To-End Autonomous Driving Framework by Vision-Language Instructed Action GenerationHaoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui 等ICCV 2025 · 被引用 19 次
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar 等ICLR 2024 · 被引用 143 次
