Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, Ivan Laptev
Abstract
To p o l o g i c a l Mapping Global Action Planning Instruction: "go into the living room and water the plant on the table." h Shortest Route Planning Next Location Local Actions Panorama Encoding Graph Update Instruction Dynamic Fusion step 𝑡+1: panorama + GPS location Coarse-scale Encoding Fine-scale Encoding a b c d e f h i g e h step 𝑡: panorama + GPS location map 𝑡-1 map 𝑡 g h j g g Figure 1 . An agent is required to navigate in unseen environments to reach target locations according to language instructions. It only obtains local observations of the environment and is allowed to make local actions, i.e., moving to neighboring locations. In this work, we propose to build topological maps on-the-fly to enable long-term action planning. The map contains visited nodes and navigable nodes that can be reached from the previously visited nodes. Our method predicts global actions, i.e., all navigable nodes in the map, and trades off complexity by combining a coarse-scale graph encoding with a fine-scale encoding of observations at the current node .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07802995-866d-4f9f-939a-5705b43dd16cCited by top-tier papers94
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.NeurIPS 2022 · 173 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang et al.ICCV 2023 · 132 citations
Builds on17
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Airbert: In-domain Pretraining for Vision-and-Language NavigationPierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev et al.ICCV 2021 · 185 citations
Related papers
- Topological Planning With Transformers for Vision-and-Language NavigationKevin Chen, Junshen K. Chen, Jo Chuang, Marynel Vázquez et al.CVPR 2021
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language NavigationZhiwei Deng, Karthik Narasimhan, Olga RussakovskyNeurIPS 2020 · 111 citations
- MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language NavigationJiaqi Chen, Bingqian Lin, Ran Xu, Zhenhua Chai et al.ACL 2024 · 29 citations
- Target-Driven Structured Transformer Planner for Vision-Language NavigationYusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang et al.ACM MM 2022 · 49 citations
- Loc4Plan: Locating Before Planning for Outdoor Vision and Language NavigationHuilin Tian, Jingke Meng, Wei-Shi Zheng, Yuan-Ming Li et al.ACM MM 2024 · 6 citations
