Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven Navigation
Hongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu, Hao Dong
摘要
The task of Visual Object Navigation (VON) involves an agent's ability to locate a particular object within a given scene. To successfully accomplish the VON task, two essential conditions must be fulfiled: 1) the user knows the name of the desired object; and 2) the user-specified object actually is present within the scene. To meet these conditions, a simulator can incorporate predefined object names and positions into the metadata of the scene. However, in real-world scenarios, it is often challenging to ensure that these conditions are always met. Humans in an unfamiliar environment may not know which objects are present in the scene, or they may mistakenly specify an object that is not actually present. Nevertheless, despite these challenges, humans may still have a demand for an object, which could potentially be fulfilled by other objects present within the scene in an equivalent manner. Hence, this paper proposes Demand-driven Navigation (DDN), which leverages the user's demand as the task instruction and prompts the agent to find an object which matches the specified demand. DDN aims to relax the stringent conditions of VON by focusing on fulfilling the user's demand rather than relying solely on specified object names. This paper proposes a method of acquiring textual attribute features of objects by extracting common sense knowledge from a large language model (LLM). These textual attribute features are subsequently aligned with visual attribute features using Contrastive Language-Image Pre-training (CLIP). Incorporating the visual attribute features as prior knowledge, enhances the navigation process. Experiments on AI2Thor with the ProcThor dataset demonstrate that the visual attribute features improve the agent's navigation performance and outperform the baseline methods commonly used in the VON and VLN task and methods with LLMs. The codes and demonstrations can be viewed at https://sites.google.com/view/demand-driven-navigation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu 等NeurIPS 2024 · 被引用 31 次
- NavBench: Probing Multimodal Large Language Models for Embodied NavigationYanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An 等NeurIPS 2025 · 被引用 27 次
- MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven NavigationHongcheng Wang, Peiqi Liu, Wenzhe Cai, Mingdong Wu 等NeurIPS 2024 · 被引用 12 次
- CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?Siqi Wang, Chao Liang, Yunfan Gao, Erxin Yu 等ICLR 2026 · 被引用 8 次
- N2M: Bridging Navigation and Manipulation by Learning Pose Preference from RolloutKaixin Chai, Hyunjun Lee, Joseph LimICML 2026 · 被引用 3 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process ThinkingYuehao Huang, Liang Liu, Shuangming Lei, Yukai Ma 等ACM MM 2025 · 被引用 1 次
- ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsBingqian Lin, Yi Zhu, Zicong Chen, Xiwen Liang 等CVPR 2022 · 被引用 45 次
- Actional Atomic-Concept Learning for Demystifying Vision-Language NavigationBingqian Lin, Yi Zhu, Xiaodan Liang, Liang Lin 等AAAI 2023 · 被引用 7 次
- Simple but Effective: CLIP Embeddings for Embodied AIApoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, Aniruddha KembhaviCVPR 2022 · 被引用 149 次
- Visual-Language Navigation Pretraining via Prompt-based Environmental Self-explorationXiwen Liang, Fengda Zhu, Lingling Li, Hang Xu 等ACL 2022 · 被引用 37 次
