Iterative Vision-and-Language Navigation
Jacob Krantz, Shurjo Banerjee, Wang Zhu, Jason J. Corso, Peter Anderson, Stefan Lee, Jesse Thomason
Abstract
We present Iterative Vision-and-Language Navigation (IVLN), a paradigm for evaluating language-guided agents navigating in a persistent environment over time. Existing Vision-and-Language Navigation (VLN) benchmarks erase the agent's memory at the beginning of every episode, testing the ability to perform cold-start navigation with no prior information. However, deployed robots occupy the same environment for long periods of time. The IVLN paradigm addresses this disparity by training and evaluating VLN agents that maintain memory across tours of scenes that consist of up to 100 ordered instruction-following Roomto-Room (R2R) episodes, each defined by an individual language instruction and a target path. We present discrete and continuous Iterative Room-to-Room (IR2R) benchmarks comprising about 400 tours each in 80 indoor scenes. We find that extending the implicit memory of high-performing transformer VLN agents is not sufficient for IVLN, but agents that build maps can benefit from environment persistence, motivating a renewed focus on map-building agents in VLN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12813d58-7c15-4fe9-b2a2-7d95a8499482Cited by top-tier papers15
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- OctoNav: Towards Generalist Embodied NavigationChen Gao, Liankai Jin, Xingyu Peng, Jiazhao Zhang et al.CVPR 2026 · 42 citations
- Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language NavigationXiang Fang, Wanlong Fang, Changshuo WangNeurIPS 2025 · 24 citations
- Towards Physically Executable 3D Gaussian for Embodied NavigationBingchen Miao, Rong Wei, Zhiqi Ge, Xiaoquan sun et al.ICLR 2026 · 22 citations
- CityNav: A Large-Scale Dataset for Real-World Aerial NavigationJungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto et al.ICCV 2025 · 8 citations
Builds on15
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- MultiON: Benchmarking Semantic Map Memory using Multi-Object NavigationSaim Wani, Shivansh Patel, Unnat Jain, Angel X. Chang et al.NeurIPS 2020 · 156 citations
- Waypoint Models for Instruction-guided Navigation in Continuous EnvironmentsJacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee et al.ICCV 2021 · 153 citations
Related papers
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
- SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language NavigationAbhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee et al.NeurIPS 2021 · 88 citations
- OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd RepresentationGanlong Zhao, Guanbin Li, Weikai Chen, Yizhou YuCVPR 2024 · 6 citations
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- Structured Scene Memory for Vision-Language NavigationHanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong et al.CVPR 2021
