Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision
Keji He, Yan Huang, Qi Wu, Jianhua Yang, Dong An, Shuanglin Sima, Liang Wang
Abstract
In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper, we address the cross-modal alignment challenge from the perspective of fine-grain. Firstly, to alleviate weak cross-modal alignment supervision from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset, namely Landmark-RxR. Secondly, to further enhance local cross-modal alignment under fine-grained supervision, we investigate the focal-oriented rewards with soft and hard forms, by focusing on the critical points sampled from fine-grained Landmark-RxR. Moreover, to fully evaluate the navigation process, we also propose a re-initialization mechanism that makes metrics insensitive to difficult points, which can cause the agent to deviate from the correct trajectories. Experimental results show that our agent has superior navigation performance on Landmark-RxR, en-RxR and R2R. Our dataset and code are available at https://github.com/hekj/Landmark-RxR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecc91685-e2a2-4af2-a132-cf014314accaCited by top-tier papers15
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- March in Chat: Interactive Prompting for Remote Embodied Referring ExpressionYanyuan Qiao, Yuankai Qi, Zheng Yu, Jing Liu et al.ICCV 2023 · 52 citations
- Frequency-Enhanced Data Augmentation for Vision-and-Language NavigationKeji He, Chenyang Si, Zhihe Lu, Yan Huang et al.NeurIPS 2023 · 32 citations
- Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationYibo Cui, Liang Xie, Yakun Zhang, Meishan Zhang et al.ICCV 2023 · 31 citations
- CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy EnvironmentsXiulong Liu, Sudipta Paul, Moitreya Chatterjee, Anoop CherianAAAI 2024 · 16 citations
Builds on7
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku et al.ICCV 2019 · 93 citations
- Neighbor-view Enhanced Model for Vision and Language NavigationDong An, Yuankai Qi, Yan Huang, Qi Wu et al.ACM MM 2021 · 71 citations
- BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby StepsWang Zhu, Hexiang Hu, Jiacheng Chen, Zhiwei Deng et al.ACL 2020 · 62 citations
- Sub-Instruction Aware Vision-and-Language NavigationYicong Hong, Cristian Rodriguez Opazo, Qi Wu, Stephen GouldEMNLP 2020 · 55 citations
Related papers
- Learning Fine-Grained Alignment for Aerial Vision-Dialog NavigationYifei Su, Dong An, Kehan Chen, Weichen Yu et al.AAAI 2025 · 7 citations
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang et al.ICCV 2023 · 132 citations
- Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success RoutesChongyang Zhao, Yuankai Qi, Qi WuACM MM 2023 · 17 citations
- CityNav: A Large-Scale Dataset for Real-World Aerial NavigationJungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto et al.ICCV 2025 · 8 citations
