DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
Wencheng Han, Dongqian Guo, Cheng-Zhong Xu, Jianbing Shen
Abstract
In the field of autonomous driving, two important features of autonomous driving car systems are the explainability of decision logic and the accuracy of environmental perception. This paper introduces DME-Driver, a new autonomous driving system that enhances the performance and reliability of autonomous driving system. DME-Driver utilizes a powerful vision language model as the decisionmaker and a planning-oriented perception model as the control signal generator. To ensure explainable and reliable driving decisions, the logical decision-maker is constructed based on a large vision language model. This model follows the logic employed by experienced human drivers and makes decisions in a similar manner. On the other hand, the generation of accurate control signals relies on precise and detailed environmental perception, which is where 3D scene perception models excel. Therefore, a planning oriented perception model is employed as the signal generator. It translates the logical decisions made by the decision-maker into accurate control signals for the self-driving cars. To effectively train the proposed model, a new dataset for autonomous driving was created. This dataset encompasses a diverse range of human driver behaviors and their underlying motivations. By leveraging this dataset, our model achieves high-precision planning accuracy through a logical thinking process. Recently, deep learning-based methods have achieved remarkable success in the realm of autonomous driving [13, 17, 28, 34, 39] . Some works [11, 24, 20, 26] proposed planning-oriented autonomous driving systems that can be trained end-to-end. As illustrated in Fig. 1 (a), this system [11] encompasses several critical perception modules, including tracking, mapping, motion, and occupancy detection. The outputs from these modules are fed into a planner, which then generates the control signals for the vehicle. This approach leverages the full potential of perception * Corresponding author Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56ac5d05-4268-4fd7-9ac4-983bc43599baCited by top-tier papers15
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li et al.ICLR 2026 · 93 citations
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous DrivingYifang Xu, Jiahao Cui, Zhihao Zhu, Hanlin Shang et al.CVPR 2026 · 18 citations
- 3D Question Answering with Scene Graph ReasoningZizhao Wu, Haohan Li, Gongyi Chen, Zhou Yu et al.ACM MM 2024 · 6 citations
- Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous DrivingXuan Chen, Shiwei Feng, Zikang Xiong, Shengwei An et al.NeurIPS 2025 · 6 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselinePenghao Wu, Xiaosong Jia, Li Chen, Junchi Yan et al.NeurIPS 2022 · 444 citations
Related papers
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous DrivingYuhang Lu, Jiadong Tu, Yuexin Ma, Xinge ZhuICCV 2025 · 1 citation
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin et al.ICCV 2025 · 8 citations
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang et al.CVPR 2026 · 36 citations
- Driving with Advice: Large Model as Motion Advisor for Joint PlanningJunyin Wang, Jinlei Yu, Hao Lin, Huikai Liu et al.AAAI 2026
- VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous DrivingFanjie Kong, Yitong Li, Weihuang Chen, Chen Min et al.ICCV 2025 · 2 citations
