VTNet: Visual Transformer Network for Object Goal Navigation
Heming Du, Xin Yu, Liang Zheng
Abstract
Object goal navigation aims to steer an agent towards a target object based on observations of the agent. It is of pivotal importance to design effective visual representations of the observed scene in determining navigation actions. In this paper, we introduce a Visual Transformer Network (VTNet) for learning informative visual representation in navigation. VTNet is a highly effective structure that embodies two key properties for visual representations: First, the relationships among all the object instances in a scene are exploited; Second, the spatial locations of objects and image regions are emphasized so that directional navigation signals can be learned. Furthermore, we also develop a pre-training scheme to associate the visual representations with navigation signals, and thus facilitate navigation policy learning. In a nutshell, VTNet embeds object and region features with their location cues as spatial-aware descriptors and then incorporates all the encoded descriptors through attention operations to achieve informative representation for navigation. Given such visual representations, agents are able to explore the correlations between visual observations and navigation actions. For example, an agent would prioritize "turning right" over "turning left" when the visual representation emphasizes on the right side of activation map. Experiments in the artificial environment AI2-Thor demonstrate that VTNet significantly outperforms state-of-the-art methods in unseen testing environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61c834a7-936b-4936-94de-c8d39c3e583aCited by top-tier papers23
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk et al.ICLR 2022 · 189 citations
- Towards Versatile Embodied NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangNeurIPS 2022 · 48 citations
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma et al.ICLR 2024 · 43 citations
- Search for or Navigate to? Dual Adaptive Thinking for Object NavigationRonghao Dang, Liuyi Wang, Zongtao He, Shuai Su et al.ICCV 2023 · 36 citations
- Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven NavigationHongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu et al.NeurIPS 2023 · 33 citations
Builds on6
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen et al.EMNLP 2020 · 158 citations
- Bayesian Relational Memory for Semantic Visual NavigationYi Wu, Yuxin Wu, Aviv Tamar, Stuart Russell et al.ICCV 2019 · 114 citations
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language NavigationZhiwei Deng, Karthik Narasimhan, Olga RussakovskyNeurIPS 2020 · 111 citations
- Situational Fusion of Visual Representation for Visual NavigationWilliam B. Shen, Danfei Xu, Yuke Zhu, Li Fei-Fei et al.ICCV 2019 · 70 citations
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin et al.CVPR 2020
Related papers
- Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation StatesHeming Du, Lincheng Li, Zi Huang, Xin YuCVPR 2023
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt et al.ICCV 2023 · 56 citations
- Unbiased Directed Object Attention Graph for Object NavigationRonghao Dang, Zhuofan Shi, Liuyi Wang, Zongtao He et al.ACM MM 2022 · 35 citations
- Hierarchical Object-to-Zone Graph for Object NavigationSixian Zhang, Xinhang Song, Yubing Bai, Weijie Li et al.ICCV 2021 · 98 citations
- ION: Instance-level Object NavigationWeijie Li, Xinhang Song, Yubing Bai, Sixian Zhang et al.ACM MM 2021 · 25 citations
