ION: Instance-level Object Navigation
Weijie Li, Xinhang Song, Yubing Bai, Sixian Zhang, Shuqiang Jiang
Abstract
Visual object navigation is a fundamental task in Embodied AI. Previous works focus on the category-wise navigation, in which navigating to any possible instance of target object category is considered a success. Those methods may be effective to find the general objects. However, it may be more practical to navigate to the specific instance in our real life, since our particular requirements are usually satisfied with specific instances rather than all instances of one category. How to navigate to the specific instance has been rarely researched before and is typically challenging to current works. In this paper, we introduce a new task of Instance Object Navigation (ION), where instance-level descriptions of targets are provided and instance-level navigation is required. In particular, multiple types of attributes such as colors, materials and object references are involved in the instance-level descriptions of the targets. In order to allow the agent to maintain the ability of instance navigation, we propose a cascade framework with Instance-Relation Graph (IRG) based navigator and instance grounding module. To specify the different instances of the same object categories, we construct instance-level graph instead of category-level one, where instances are regarded as nodes, encoded with the representation of colors, materials and locations (bounding boxes). During navigation, the detected instances can activate corresponding nodes in IRG, which are updated with graph convolutional neural network (GCNN). The final instance prediction is obtained with the grounding module by selecting the candidates (instances) with maximum probability (a joint probability of category, color and material, obtained by corresponding regressors with softmax). For the task evaluation, we build a benchmark for instance-level object navigation on AI2-Thor simulator, where over 27,735 object instance descriptions and navigation groundtruth are automatically obtained through the interaction with the simulator. The proposed model outperforms the baseline in instance-level metrics, showing that our proposed graph model can guide instance object navigation, as well as leaving promising room for further improvement. The project is available at https://github.com/LWJ312/ION.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f15efcb3-9424-4609-af02-442cf2e635e4Cited by top-tier papers9
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
- Unbiased Directed Object Attention Graph for Object NavigationRonghao Dang, Zhuofan Shi, Liuyi Wang, Zongtao He et al.ACM MM 2022 · 35 citations
- Modeling Dynamic Environments with Scene Graph MemoryAndrey Kurenkov, Michael Lingelbach, Tanmay Agarwal, Emily Jin et al.ICML 2023 · 21 citations
- Implicit Obstacle Map-driven Indoor Navigation Model for Robust Obstacle AvoidanceWei Xie, Haobo Jiang, Shuo Gu, Jin XieACM MM 2023 · 9 citations
- NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object NavigationWei Xie, Haobo Jiang, Yun Zhu, Jianjun Qian et al.AAAI 2025 · 7 citations
Related papers
- Hierarchical Object-to-Zone Graph for Object NavigationSixian Zhang, Xinhang Song, Yubing Bai, Weijie Li et al.ICCV 2021 · 98 citations
- Instance-Aware Exploration-Verification-Exploitation for Instance ImageGoal NavigationXiaohan Lei, Min Wang, Wengang Zhou, Li Li et al.CVPR 2024
- VTNet: Visual Transformer Network for Object Goal NavigationHeming Du, Xin Yu, Liang ZhengICLR 2021 · 36 citations
- Navigating to Objects Specified by ImagesJacob Krantz, Théophile Gervet, Karmesh Yadav, Austin S. Wang et al.ICCV 2023 · 70 citations
- Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationBolei Chen, Jiaxu Kang, Ping Zhong, Yixiong Liang et al.ACM MM 2024 · 4 citations
