VTNet: Visual Transformer Network for Object Goal Navigation
Heming Du, Xin Yu, Liang Zheng
摘要
Object goal navigation aims to steer an agent towards a target object based on observations of the agent. It is of pivotal importance to design effective visual representations of the observed scene in determining navigation actions. In this paper, we introduce a Visual Transformer Network (VTNet) for learning informative visual representation in navigation. VTNet is a highly effective structure that embodies two key properties for visual representations: First, the relationships among all the object instances in a scene are exploited; Second, the spatial locations of objects and image regions are emphasized so that directional navigation signals can be learned. Furthermore, we also develop a pre-training scheme to associate the visual representations with navigation signals, and thus facilitate navigation policy learning. In a nutshell, VTNet embeds object and region features with their location cues as spatial-aware descriptors and then incorporates all the encoded descriptors through attention operations to achieve informative representation for navigation. Given such visual representations, agents are able to explore the correlations between visual observations and navigation actions. For example, an agent would prioritize "turning right" over "turning left" when the visual representation emphasizes on the right side of activation map. Experiments in the artificial environment AI2-Thor demonstrate that VTNet significantly outperforms state-of-the-art methods in unseen testing environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk 等ICLR 2022 · 被引用 189 次
- Towards Versatile Embodied NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangNeurIPS 2022 · 被引用 48 次
- GROOT: Learning to Follow Instructions by Watching Gameplay VideosShaofei Cai, Bowei Zhang, Zihao Wang, Xiaojian Ma 等ICLR 2024 · 被引用 43 次
- Search for or Navigate to? Dual Adaptive Thinking for Object NavigationRonghao Dang, Liuyi Wang, Zongtao He, Shuai Su 等ICCV 2023 · 被引用 36 次
- Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven NavigationHongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu 等NeurIPS 2023 · 被引用 33 次
它引用的顶会 Paper6
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen 等EMNLP 2020 · 被引用 158 次
- Bayesian Relational Memory for Semantic Visual NavigationYi Wu, Yuxin Wu, Aviv Tamar, Stuart Russell 等ICCV 2019 · 被引用 114 次
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language NavigationZhiwei Deng, Karthik Narasimhan, Olga RussakovskyNeurIPS 2020 · 被引用 111 次
- Situational Fusion of Visual Representation for Visual NavigationWilliam B. Shen, Danfei Xu, Yuke Zhu, Li Fei-Fei 等ICCV 2019 · 被引用 70 次
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等CVPR 2020
相关 Paper
- Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation StatesHeming Du, Lincheng Li, Zi Huang, Xin YuCVPR 2023
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt 等ICCV 2023 · 被引用 56 次
- Unbiased Directed Object Attention Graph for Object NavigationRonghao Dang, Zhuofan Shi, Liuyi Wang, Zongtao He 等ACM MM 2022 · 被引用 35 次
- Hierarchical Object-to-Zone Graph for Object NavigationSixian Zhang, Xinhang Song, Yubing Bai, Weijie Li 等ICCV 2021 · 被引用 98 次
- ION: Instance-level Object NavigationWeijie Li, Xinhang Song, Yubing Bai, Sixian Zhang 等ACM MM 2021 · 被引用 25 次
