MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
Lingjun Zhang, Yujian Yuan, Changjie Wu, Xinyuan Chang, Xin Cai, Shuang Zeng, Linzhe Shi, Sijin Wang, Hang Zhang, Mu Xu
Abstract
Vision-Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end-to-end autonomous driving systems. Chain-of-Thought (CoT), as VLM's widely used reasoning strategy, is facing critical challenges. Existing textual CoT has a large gap between text semantic space and trajectory physical space. Although the recent approach utilizes future image to replace text as CoT process, it lacks clear planning-oriented objective guidance to generate images with accurate scene evolution. To address these, we innovatively propose MindDriver, a progressive multimodal reasoning framework that enables VLM to imitate human-like progressive thinking for autonomous driving. MindDriver presents semantic understanding, semantic-to-physical space imagination, and physical-space trajectory planning. To achieve aligned reasoning processes in MindDriver, we develop a feedback-guided automatic data annotation pipeline to generate aligned multimodal reasoning training data. Furthermore, we develop a progressive reinforcement fine-tuning method to optimize the alignment through progressive high- level reward-based learning. MindDriver demonstrates superior performance in both nuScences open-loop and Bench2Drive closed-loop evaluation. Codes are available at https://github.com/hotdogcheesewhite/MindDriver.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc7bac37-a3ec-4ad3-ac75-b4af47dff80aCited by top-tier papers1
Ask how each one uses itBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselinePenghao Wu, Xiaosong Jia, Li Chen, Junchi Yan et al.NeurIPS 2022 · 444 citations
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
Related papers
- Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement LearningYue Li, Meng Tian, Dechang Zhu, Jiangtong Zhu et al.AAAI 2026 · 27 citations
- Latent Chain-of-Thought World Modeling for End-to-End Autonomous DrivingShuhan Tan, Kashyap Chitta, Yuxiao Chen, Ran Tian et al.CVPR 2026 · 10 citations
- AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous DrivingZhenlong Yuan, Chengxuan Qian, Jing Tang, Rui Chen et al.ICLR 2026 · 24 citations
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingShuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu et al.NeurIPS 2025 · 228 citations
- AutoDrive-P3: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-TuningYuqi Ye, Zijian Zhang, Junhong Lin, Shangkun Sun et al.ICLR 2026 · 16 citations
