Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring Expression
Chen Gao, Jinyu Chen, Si Liu, Luting Wang, Qiong Zhang, Qi Wu
摘要
The Remote Embodied Referring Expression (REVERIE) is a recently raised task that requires an agent to navigate to and localise a referred remote object according to a high-level language instruction. Different from related VLN tasks, the key to REVERIE is to conduct goal-oriented exploration instead of strict instruction-following, due to the lack of step-by-step navigation guidance. In this paper, we propose a novel Cross-modality Knowledge Reasoning (CKR) model to address the unique challenges of this task. The CKR, based on a transformer-architecture, learns to generate scene memory tokens and utilise these informative history clues for exploration. Particularly, a Roomand-Object Aware Attention (ROAA) mechanism is devised to explicitly perceive the room-and object-type information from both linguistic and visual observations. Moreover, through incorporating commonsense knowledge, we propose a Knowledge-enabled Entity Relationship Reasoning (KERR) module to learn the internal-external correlations among room-and object-entities for agent to make proper action at each viewpoint. Evaluation on REVERIE benchmark demonstrates the superiority of the CKR model, which significantly boosts SPL and REVERIE-success rate by 64.67% and 46.05%, respectively. Code is available at: https://github.com/alloldman/CKR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- NodeFormer: A Scalable Graph Structure Learning Transformer for Node ClassificationQitian Wu, Wentao Zhao, Zenan Li, David P. Wipf 等NeurIPS 2022 · 被引用 472 次
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
- Simplifying and Empowering Transformers for Large-Graph RepresentationsQitian Wu, Wentao Zhao, Chenxiao Yang, Hengrui Zhang 等NeurIPS 2023 · 被引用 318 次
- Graph Structure Learning with Variational Information BottleneckQingyun Sun, Jianxin Li, Hao Peng, Jia Wu 等AAAI 2022 · 被引用 224 次
- 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive SelectionJunyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao 等CVPR 2022 · 被引用 72 次
它引用的顶会 Paper15
- SUGAR: Subgraph Neural Network with Reinforcement Pooling and Self-Supervised Mutual Information MechanismQingyun Sun, Jianxin Li, Hao Peng, Jia Wu 等WWW 2021 · 被引用 196 次
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu 等NeurIPS 2020 · 被引用 167 次
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language NavigationZhiwei Deng, Karthik Narasimhan, Olga RussakovskyNeurIPS 2020 · 被引用 111 次
- Transferable Representation Learning in Vision-and-Language NavigationHaoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku 等ICCV 2019 · 被引用 93 次
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 被引用 54 次
相关 Paper
- Scene-Intuitive Agent for Remote Embodied Visual GroundingXiangru Lin, Guanbin Li, Yizhou YuCVPR 2021
- Augmented Commonsense Knowledge for Remote Object GroundingBahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu 等AAAI 2024 · 被引用 21 次
- Layout-Aware Dreamer for Embodied Visual Referring Expression GroundingMingxiao Li, Zehao Wang, Tinne Tuytelaars, Marie-Francine MoensAAAI 2023 · 被引用 22 次
- KERM: Knowledge Enhanced Reasoning for Vision-and-Language NavigationXiangyang Li, Zihan Wang, Jiahao Yang, Yaowei Wang 等CVPR 2023
- March in Chat: Interactive Prompting for Remote Embodied Referring ExpressionYanyuan Qiao, Yuankai Qi, Zheng Yu, Jing Liu 等ICCV 2023 · 被引用 52 次
