Distilling Multi-modal Large Language Models for Autonomous Driving
Deepti Hegde, Rajeev Yasarla, Hong Cai, Shizhong Han, Apratim Bhattacharyya, Shweta Mahajan, Litian Liu, Risheek Garrepalli, Vishal M. Patel, Fatih Porikli
Abstract
Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational costs. To address this, we propose DiMA, an end-to-end autonomous driving system that maintains the efficiency of an LLM-free (or vision-based) planner while leveraging the world knowledge of an LLM. DiMA distills the information from a multi-modal LLM to a visionbased end-to-end planner through a set of specially designed surrogate tasks. Under a joint training strategy, a scene encoder common to both networks produces structured representations that are semantically grounded as well as aligned to the final planning objective. Notably, the LLM is optional at inference, enabling robust planning without compromising on efficiency. Training with DiMA results in a 37% reduction in the L2 trajectory error and an 80% reduction in the collision rate of the vision-based planner, as well as a 44% trajectory error reduction in longtail scenarios. DiMA also achieves state-of-the-art performance on the nuScenes planning benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c695223b-d98f-4bd9-8ba8-67bcaf648f34Cited by top-tier papers6
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous DrivingZhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li et al.CVPR 2026 · 108 citations
- RoCA: Robust Cross-Domain End-to-End Autonomous DrivingRajeev Yasarla, Shizhong Han, Hsin-Pai Cheng, Apratim Bhattacharyya et al.ICML 2026 · 8 citations
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous DrivingJiahao Wang, Bo Sun, Yijing Bai, Vincent Casser et al.CVPR 2026 · 2 citations
- RobusTor3D: Robust Multimodal 3D Object Detector for Autonomous Driving by Vision-Language Knowledge BlendingYing Yang, Hui Yin, Aixin Chong, Hui Wang et al.AAAI 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
Related papers
- VLP: Vision Language Planning for Autonomous DrivingChenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik et al.CVPR 2024 · 53 citations
- VLMPlanner: Integrating Visual Language Models with Motion PlanningZhipeng Tang, Sha Zhang, Jiajun Deng, Chenjie Wang et al.ACM MM 2025 · 2 citations
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationYichen Xie, Runsheng Xu, Tong He, Jyh-Jing Hwang et al.CVPR 2025
- Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous DrivingTengpeng Li, Hanli Wang, Xianfei Li, Wenlong Liao et al.AAAI 2025 · 16 citations
