SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language Navigation
Abhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee, Dhruv Batra
Abstract
Natural language instructions for visual navigation often use scene descriptions (e.g., 'bedroom') and object references (e.g., 'green chairs') to provide a breadcrumb trail to a goal location. This work presents a transformer-based vision-andlanguage navigation (VLN) agent that uses two different visual encoders -a scene classification network and an object detector -which produce features that match these two distinct types of visual cues. In our method, scene features contribute high-level contextual information that supports object-level processing. With this design, our model is able to use vision-and-language pretraining (i.e., learning the alignment between images and text from large-scale web data) to substantially improve performance on the Room-to-Room (R2R) [1] and Room-Across-Room (RxR) [2] benchmarks. Specifically, our approach leads to improvements of 1.8% absolute in SPL on R2R and 3.7% absolute in SR on RxR. Our analysis reveals even larger gains for navigation instructions that contain six or more object references, which further suggests that our approach is better able to use object features and align them to references in the instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71beb170-c4e4-412c-aed5-b469ae5492d1Cited by top-tier papers25
- Bird's-Eye-View Scene Graph for Vision-Language NavigationRui Liu, Xiaohan Wang, Wenguan Wang, Yi YangICCV 2023 · 100 citations
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 76 citations
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li et al.ICCV 2023 · 57 citations
- Target-Driven Structured Transformer Planner for Vision-Language NavigationYusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang et al.ACM MM 2022 · 49 citations
- Unbiased Directed Object Attention Graph for Object NavigationRonghao Dang, Zhuofan Shi, Liuyi Wang, Zongtao He et al.ACM MM 2022 · 35 citations
Builds on5
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu et al.NeurIPS 2020 · 167 citations
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang et al.CVPR 2020
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- Vision-Language Navigation With Self-Supervised Auxiliary Reasoning TasksFengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan LiangCVPR 2020
Related papers
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku et al.ICCV 2019 · 93 citations
- Neighbor-view Enhanced Model for Vision and Language NavigationDong An, Yuankai Qi, Yan Huang, Qi Wu et al.ACM MM 2021 · 71 citations
- Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationYibo Cui, Liang Xie, Yakun Zhang, Meishan Zhang et al.ICCV 2023 · 31 citations
- A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation LearningAishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh et al.CVPR 2023
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
