Language and Visual Entity Relationship Graph for Agent Navigation
Yicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu, Stephen Gould
Abstract
Vision-and-Language Navigation (VLN) requires an agent to navigate in a real-world environment following natural language instructions. From both the textual and visual perspectives, we find that the relationships among the scene, its objects,and directional clues are essential for the agent to interpret complex instructions and correctly perceive the environment. To capture and utilize the relationships, we propose a novel Language and Visual Entity Relationship Graph for modelling the inter-modal relationships between text and vision, and the intra-modal relationships among visual entities. We propose a message passing algorithm for propagating information between language elements and visual entities in the graph, which we then combine to determine the next action to take. Experiments show that by taking advantage of the relationships we are able to improve over state-of-the-art. On the Room-to-Room (R2R) benchmark, our method achieves the new best performance on the test unseen split with success rate weighted by path length (SPL) of 52%. On the Room-for-Room (R4R) dataset, our method significantly improves the previous best from 13% to 34% on the success weighted by normalized dynamic time warping (SDTW). Code is available at: this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c8d7d57-d568-4447-ae4b-695c050d6034Cited by top-tier papers52
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
Builds on5
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
- Sub-Instruction Aware Vision-and-Language NavigationYicong Hong, Cristian Rodriguez Opazo, Qi Wu, Stephen GouldEMNLP 2020 · 55 citations
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang et al.CVPR 2020
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
- Vision-Language Navigation With Self-Supervised Auxiliary Reasoning TasksFengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan LiangCVPR 2020
Related papers
- SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language NavigationAbhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee et al.NeurIPS 2021 · 88 citations
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
- Neighbor-view Enhanced Model for Vision and Language NavigationDong An, Yuankai Qi, Yan Huang, Qi Wu et al.ACM MM 2021 · 71 citations
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 29 citations
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku et al.ICCV 2019 · 93 citations
