VLN-Trans: Translator for the Vision and Language Navigation Agent
Yue Zhang, Parisa Kordjamshidi
Abstract
Language understanding is essential for the navigation agent to follow instructions. We observe two kinds of issues in the instructions that can make the navigation task challenging: 1. The mentioned landmarks are not recognizable by the navigation agent due to the different vision abilities of the instructor and the modeled agent. 2. The mentioned landmarks are applicable to multiple targets, thus not distinctive for selecting the target among the candidate viewpoints. To deal with these issues, we design a translator module for the navigation agent to convert the original instructions into easy-tofollow sub-instruction representations at each step. The translator needs to focus on the recognizable and distinctive landmarks based on the agent's visual abilities and the observed visual environment. To achieve this goal, we create a new synthetic sub-instruction dataset and design specific tasks to train the translator and the navigation agent. We evaluate our approach on Room2Room (R2R), Room4room (R4R), and Room2Room Last (R2R-Last) datasets and achieve state-of-the-art results on multiple benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37ca743e-cd5c-4937-8090-780f7965ea29Cited by top-tier papers11
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and NavigationZihan Wang, Seungjun Lee, Guangzhao Dai, Gim Hee LeeCVPR 2026 · 9 citations
- AeroDuo: Aerial Duo for UAV-based Vision and Language NavigationRuipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang et al.ACM MM 2025 · 6 citations
- NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional GeneralizationDanial Kamali, Elham J. Barezi, Parisa KordjamshidiAAAI 2025 · 4 citations
- Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation AgentsTianyi Ma, Yue Zhang, Zehao Wang, Parisa KordjamshidiACL 2026 · 3 citations
- Narrowing the Gap between Vision and Action in NavigationYue Zhang, Parisa KordjamshidiACM MM 2024 · 2 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu et al.NeurIPS 2020 · 167 citations
Related papers
- Sub-Instruction Aware Vision-and-Language NavigationYicong Hong, Cristian Rodriguez Opazo, Qi Wu, Stephen GouldEMNLP 2020 · 55 citations
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
- Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationYibo Cui, Liang Xie, Yakun Zhang, Meishan Zhang et al.ICCV 2023 · 31 citations
- Less is More: Generating Grounded Navigation Instructions from LandmarksSu Wang, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar et al.CVPR 2022 · 41 citations
- A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation LearningAishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh et al.CVPR 2023
