DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Driving
Dongqian Guo, Haoran Wei, Wencheng Han, Runzhou Tao, Zhongying Qiu, Jianfei Yang, Jianbing Shen
Abstract
Autonomous driving has made substantial progress recently, achieving reliable performance in most real-world environments. However, existing algorithms still depend heavily on high-definition maps, making them ineffective in mapless scenarios such as indoor parking lots. These limitations hinder seamless point-to-point navigation and restrict the broader deployment of the autonomous driving system. To address this challenge, we propose DriveVLN, a new task that extends Vision-and-Language Navigation (VLN) to autonomous driving. DriveVLN employs visual and linguistic priors to guide vehicles toward destinations based solely on concise natural-language descriptions, without access to predefined maps or routes. Unlike conventional VLN, which relies on detailed step-wise instructions in indoor environments, DriveVLN requires models to produce navigation information based on diverse visual cues and history, including signs, landmarks, and textual indicators. We further develop a CARLA-based simulation engine comprising over 200 realistic scenes reconstructed from real road scans, enabling large-scale training and closed-loop evaluation. A baseline model is established through supervised fine-tuning on real data, followed by reinforcement learning in simulation. Comprehensive experiments show that DriveVLN effectively bridges map-based and mapless driving, providing a new foundation for unified, language-driven autonomous navigation in complex real-world environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5f882aa-6896-4a1f-a743-4b8b9edf4df2Builds on14
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao et al.AAAI 2024 · 314 citations
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic PlanningBo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao et al.ICLR 2026 · 259 citations
- LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsHao Shao, Yuxuan Hu, Letian Wang, Guanglu Song et al.CVPR 2024 · 114 citations
- Learning to Follow Directions in Street ViewKarl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath et al.AAAI 2020 · 78 citations
- DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous DrivingWencheng Han, Dongqian Guo, Cheng-Zhong Xu, Jianbing ShenAAAI 2025 · 64 citations
Related papers
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li et al.CVPR 2026 · 7 citations
- VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationJialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit BansalAAAI 2024 · 13 citations
- MapDream: Task-Driven Map Learning for Vision-Language NavigationGuoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang et al.ICML 2026
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 29 citations
- IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor EnvironmentsXu Liu, Yu Liu, Hanshuo Qiu, Qirong Yang et al.AAAI 2026 · 5 citations
