VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-Grounded Autonomous Driving
Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang, Xiang Wan, Xiaonan Luo, Guanbin Li
Abstract
Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLMs). However, current LLM-based approaches encounter critical challenges: (1) Failure analysis reveals that frequent collisions and obstructions, stemming from limitations in visual representations, remain primary obstacles to robust driving performance. (2) The substantial parameters of LLMs pose considerable deployment hurdles. To address these limitations, we introduce VLDrive, a novel approach featuring a lightweight MLLM architecture with enhanced vision components. VLDrive achieves compact visual tokens through innovative strategies, including cycle-consistent dynamic visual pruning and memory-enhanced feature aggregation. Furthermore, we propose a distance-decoupled instruction attention mechanism to improve joint visual-linguistic feature learning, particularly for long-range visual tokens. Extensive experiments conducted in the CARLA simulator demonstrate VLDrive's effectiveness. Notably, VLDrive achieves state-of-the-art driving performance while reducing parameters by 81% (from 7B to 1.3B), yielding substantial driving score improvements of , and at tiny, short, and long distances, respectively, in closed-loop evaluations. Code is available at https://github.com/ReaFly/VLDrive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention DiscrepancyYutong Xie, Zhenglin Hua, Ran Wang, Wing W. Y. Ng et al.ICML 2026 · 1 citation
- Rethinking Instruction Drift as a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic PlanningKewei Chen, Yayu Long, mingsheng shangICML 2026
- StreamRAG: Enhancing Real-Time Video Understanding with Retrieval AugmentationJunlin Xie, Quanlong Zheng, Ruifei Zhang, Kuo Wang et al.CVPR 2026
Builds on17
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
Related papers
- Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous DrivingMinhao Xiong, Zichen Wen, Zhuangcheng Gu, Xuyang Liu et al.CVPR 2026 · 17 citations
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang et al.CVPR 2026 · 36 citations
- Ask Less, See More: Communication-Conditioned Token Pruning for Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language ModelsShiqi Sun, Yantao Lu, Bingkun Sun, Ning Liu et al.ICML 2026
- LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsHao Shao, Yuxuan Hu, Letian Wang, Guanglu Song et al.CVPR 2024 · 114 citations
- Growing a Twig to Accelerate Large Vision-Language ModelsZhenwei Shao, Mingyang Wang, Zhou Yu, Wenwen Pan et al.ICCV 2025 · 3 citations
