A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning
Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, Zarana Parekh
Abstract
Recent studies in Vision-and-Language Navigation (VLN) train RL agents to execute natural-language navigation instructions in photorealistic environments, as a step towards robots that can follow human instructions. However, given the scarcity of human instruction data and limited diversity in the training environments, these agents still struggle with complex language grounding and spatial language understanding. Pretraining on large text and image-text datasets from the web has been extensively explored but the improvements are limited. We investigate large-scale augmentation with synthetic instructions. We take 500+ indoor environments captured in denselysampled 360 • panoramas, construct navigation trajectories through these panoramas, and generate a visuallygrounded instruction for each trajectory using Marky [63], a high-quality multilingual navigation instruction generator. We also synthesize image observations from novel viewpoints using an image-to-image GAN [27]. The resulting dataset of 4.2M instruction-trajectory pairs is two orders of magnitude larger than existing human-annotated datasets, and contains a wider variety of environments and viewpoints. To efficiently leverage data at this scale, we train a simple transformer agent with imitation learning. On the challenging RxR dataset, our approach outperforms all existing RL agents, improving the state-of-the-art NDTW from 71.1 to 79.1 in seen environments, and from 64.6 to 66.8 in unseen test environments. Our work points to a new path to improving instruction-following agents, emphasizing largescale training on near-human quality synthetic instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf494f5e-3049-47be-b346-81396723788fCited by top-tier papers28
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
- DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous DrivingWencheng Han, Dongqian Guo, Cheng-Zhong Xu, Jianbing ShenAAAI 2025 · 64 citations
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li et al.ICCV 2023 · 57 citations
- NavBench: Probing Multimodal Large Language Models for Embodied NavigationYanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An et al.NeurIPS 2025 · 27 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationJialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit BansalAAAI 2024 · 13 citations
- PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationJialu Li, Mohit BansalNeurIPS 2023 · 110 citations
- Less is More: Generating Grounded Navigation Instructions from LandmarksSu Wang, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar et al.CVPR 2022 · 41 citations
- RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied NavigationMingfei Han, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova et al.CVPR 2025
- SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language NavigationAbhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee et al.NeurIPS 2021 · 88 citations
