VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
Jialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit Bansal
Abstract
Outdoor Vision-and-Language Navigation (VLN) requires an agent to navigate through realistic 3D outdoor environments based on natural language instructions. The performance of existing VLN methods is limited by insufficient diversity in navigation environments and limited training data. To address these issues, we propose VLN-Video, which utilizes the diverse outdoor environments present in driving videos in multiple cities in the U.S. augmented with automatically generated navigation instructions and actions to improve outdoor VLN performance. VLN-Video combines the best of intuitive classical approaches and modern deep learning techniques, using template infilling to generate grounded non-repetitive navigation instructions, combined with an image rotation similarity based navigation action predictor to obtain VLN style data from driving videos for pretraining deep learning VLN models. We pre-train the model on the Touchdown dataset and our video-augmented dataset created from driving videos with three proxy tasks: Masked Language Modeling, Instruction and Trajectory Matching, and Next Action Prediction, so as to learn temporally-aware and visually-aligned instruction representations. The learned instruction representation is adapted to the state-of-the-art navigation agent when fine-tuning on the Touchdown dataset. Empirical results demonstrate that VLN-Video significantly outperforms previous state-of-the-art models by 2.1% in task completion rate, achieving a new state-of-the-art on the Touchdown dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1089bcdc-75f9-4da8-afe3-5725e5e277c6Cited by top-tier papers5
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor NavigationXia Su, Ruiqi Chen, Benlin Liu, Jingwei Ma et al.CVPR 2026 · 8 citations
- Loc4Plan: Locating Before Planning for Outdoor Vision and Language NavigationHuilin Tian, Jingke Meng, Wei-Shi Zheng, Yuan-Ming Li et al.ACM MM 2024 · 6 citations
- FLAME: Learning to Navigate with Multimodal LLM in Urban EnvironmentsYunzhe Xu, Yiyuan Pan, Zhe Liu, Hesheng WangAAAI 2025 · 3 citations
- Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion ControlHao Ren, Zetong Bi, Yiming Zeng, Le Zheng et al.ICML 2026 · 1 citation
- FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AIYuhang Peng, Yizhou Pan, Xinning He, Jihaoyu Yang et al.AAAI 2026 · 1 citation
Builds on16
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Airbert: In-domain Pretraining for Vision-and-Language NavigationPierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev et al.ICCV 2021 · 185 citations
- Vision-Language Navigation with Random Environmental MixupChong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang et al.ICCV 2021 · 113 citations
Related papers
- A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation LearningAishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh et al.CVPR 2023
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li et al.ICCV 2023 · 57 citations
- RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied NavigationMingfei Han, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova et al.CVPR 2025
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor AreasRaphael Schumann, Stefan RiezlerACL 2022 · 38 citations
