Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring Expression
Chen Gao, Jinyu Chen, Si Liu, Luting Wang, Qiong Zhang, Qi Wu
Abstract
The Remote Embodied Referring Expression (REVERIE) is a recently raised task that requires an agent to navigate to and localise a referred remote object according to a high-level language instruction. Different from related VLN tasks, the key to REVERIE is to conduct goal-oriented exploration instead of strict instruction-following, due to the lack of step-by-step navigation guidance. In this paper, we propose a novel Cross-modality Knowledge Reasoning (CKR) model to address the unique challenges of this task. The CKR, based on a transformer-architecture, learns to generate scene memory tokens and utilise these informative history clues for exploration. Particularly, a Roomand-Object Aware Attention (ROAA) mechanism is devised to explicitly perceive the room-and object-type information from both linguistic and visual observations. Moreover, through incorporating commonsense knowledge, we propose a Knowledge-enabled Entity Relationship Reasoning (KERR) module to learn the internal-external correlations among room-and object-entities for agent to make proper action at each viewpoint. Evaluation on REVERIE benchmark demonstrates the superiority of the CKR model, which significantly boosts SPL and REVERIE-success rate by 64.67% and 46.05%, respectively. Code is available at: https://github.com/alloldman/CKR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fbede35-e12d-491e-9774-a786487a3887Cited by top-tier papers29
- NodeFormer: A Scalable Graph Structure Learning Transformer for Node ClassificationQitian Wu, Wentao Zhao, Zenan Li, David P. Wipf et al.NeurIPS 2022 · 472 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Simplifying and Empowering Transformers for Large-Graph RepresentationsQitian Wu, Wentao Zhao, Chenxiao Yang, Hengrui Zhang et al.NeurIPS 2023 · 318 citations
- Graph Structure Learning with Variational Information BottleneckQingyun Sun, Jianxin Li, Hao Peng, Jia Wu et al.AAAI 2022 · 224 citations
- 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive SelectionJunyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao et al.CVPR 2022 · 72 citations
Builds on15
- SUGAR: Subgraph Neural Network with Reinforcement Pooling and Self-Supervised Mutual Information MechanismQingyun Sun, Jianxin Li, Hao Peng, Jia Wu et al.WWW 2021 · 196 citations
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu et al.NeurIPS 2020 · 167 citations
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language NavigationZhiwei Deng, Karthik Narasimhan, Olga RussakovskyNeurIPS 2020 · 111 citations
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku et al.ICCV 2019 · 93 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
Related papers
- Scene-Intuitive Agent for Remote Embodied Visual GroundingXiangru Lin, Guanbin Li, Yizhou YuCVPR 2021
- Augmented Commonsense Knowledge for Remote Object GroundingBahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu et al.AAAI 2024 · 21 citations
- Layout-Aware Dreamer for Embodied Visual Referring Expression GroundingMingxiao Li, Zehao Wang, Tinne Tuytelaars, Marie-Francine MoensAAAI 2023 · 22 citations
- KERM: Knowledge Enhanced Reasoning for Vision-and-Language NavigationXiangyang Li, Zihan Wang, Jiahao Yang, Yaowei Wang et al.CVPR 2023
- March in Chat: Interactive Prompting for Remote Embodied Referring ExpressionYanyuan Qiao, Yuankai Qi, Zheng Yu, Jing Liu et al.ICCV 2023 · 52 citations
