ION: Instance-level Object Navigation
Weijie Li, Xinhang Song, Yubing Bai, Sixian Zhang, Shuqiang Jiang
摘要
Visual object navigation is a fundamental task in Embodied AI. Previous works focus on the category-wise navigation, in which navigating to any possible instance of target object category is considered a success. Those methods may be effective to find the general objects. However, it may be more practical to navigate to the specific instance in our real life, since our particular requirements are usually satisfied with specific instances rather than all instances of one category. How to navigate to the specific instance has been rarely researched before and is typically challenging to current works. In this paper, we introduce a new task of Instance Object Navigation (ION), where instance-level descriptions of targets are provided and instance-level navigation is required. In particular, multiple types of attributes such as colors, materials and object references are involved in the instance-level descriptions of the targets. In order to allow the agent to maintain the ability of instance navigation, we propose a cascade framework with Instance-Relation Graph (IRG) based navigator and instance grounding module. To specify the different instances of the same object categories, we construct instance-level graph instead of category-level one, where instances are regarded as nodes, encoded with the representation of colors, materials and locations (bounding boxes). During navigation, the detected instances can activate corresponding nodes in IRG, which are updated with graph convolutional neural network (GCNN). The final instance prediction is obtained with the grounding module by selecting the candidates (instances) with maximum probability (a joint probability of category, color and material, obtained by corresponding regressors with softmax). For the task evaluation, we build a benchmark for instance-level object navigation on AI2-Thor simulator, where over 27,735 object instance descriptions and navigation groundtruth are automatically obtained through the interaction with the simulator. The proposed model outperforms the baseline in instance-level metrics, showing that our proposed graph model can guide instance object navigation, as well as leaving promising room for further improvement. The project is available at https://github.com/LWJ312/ION.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper9
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等ICCV 2023 · 被引用 136 次
- Unbiased Directed Object Attention Graph for Object NavigationRonghao Dang, Zhuofan Shi, Liuyi Wang, Zongtao He 等ACM MM 2022 · 被引用 35 次
- Modeling Dynamic Environments with Scene Graph MemoryAndrey Kurenkov, Michael Lingelbach, Tanmay Agarwal, Emily Jin 等ICML 2023 · 被引用 21 次
- Implicit Obstacle Map-driven Indoor Navigation Model for Robust Obstacle AvoidanceWei Xie, Haobo Jiang, Shuo Gu, Jin XieACM MM 2023 · 被引用 9 次
- NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object NavigationWei Xie, Haobo Jiang, Yun Zhu, Jianjun Qian 等AAAI 2025 · 被引用 7 次
相关 Paper
- Hierarchical Object-to-Zone Graph for Object NavigationSixian Zhang, Xinhang Song, Yubing Bai, Weijie Li 等ICCV 2021 · 被引用 98 次
- Instance-Aware Exploration-Verification-Exploitation for Instance ImageGoal NavigationXiaohan Lei, Min Wang, Wengang Zhou, Li Li 等CVPR 2024
- VTNet: Visual Transformer Network for Object Goal NavigationHeming Du, Xin Yu, Liang ZhengICLR 2021 · 被引用 36 次
- Navigating to Objects Specified by ImagesJacob Krantz, Théophile Gervet, Karmesh Yadav, Austin S. Wang 等ICCV 2023 · 被引用 70 次
- Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationBolei Chen, Jiaxu Kang, Ping Zhong, Yixiong Liang 等ACM MM 2024 · 被引用 4 次
