Self-Motivated Communication Agent for Real-World Vision-Dialog Navigation
Yi Zhu, Yue Weng, Fengda Zhu, Xiaodan Liang, Qixiang Ye, Yutong Lu, Jianbin Jiao
Abstract
Vision-Dialog Navigation (VDN) requires an agent to ask questions and navigate following the human responses to find target objects. Conventional approaches are only allowed to ask questions at predefined locations, which are built upon expensive dialogue annotations, and inconvenience the real-word human-robot communication and cooperation. In this paper, we propose a Self-Motivated Communication Agent (SCoA) that learns whether and what to communicate with human adaptively to acquire instructive information for realizing dialogue annotation-free navigation and enhancing the transferability in real-world unseen environment. Specifically, we introduce a whether-to-ask (WeTA) policy, together with uncertainty of which action to choose, to indicate whether the agent should ask a question. Then, a what-to-ask (WaTA) policy is proposed, in which, along with the oracle’s answers, the agent learns to score question candidates so as to pick up the most informative one for navigation, and meanwhile mimic oracle’s answering. Thus, the agent can navigate in a self-Q&A manner even in real-world environment where the human assistance is often unavailable. Through joint optimization of communication and navigation in a unified imitation learning and reinforcement learning framework, SCoA asks a question if necessary and obtains a hint for guiding the agent to move towards the target with less communication cost. Experiments on seen and unseen environments demonstrate that SCoA shows not only superior performance over existing baselines without dialog annotations, but also competing results compared with rich dialog annotations based counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a426cf2-ab76-4a2f-a0d1-d894b9bcb1c1Cited by top-tier papers10
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- AVLEN: Audio-Visual-Language Embodied Navigation in 3D EnvironmentsSudipta Paul, Amit Roy-Chowdhury, Anoop CherianNeurIPS 2022 · 43 citations
- Frequency-Enhanced Data Augmentation for Vision-and-Language NavigationKeji He, Chenyang Si, Zhihe Lu, Yan Huang et al.NeurIPS 2023 · 32 citations
- LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction FollowingCheng-Fu Yang, Yen-Chun Chen, Jianwei Yang, Xiyang Dai et al.EMNLP 2023 · 6 citations
- Learn How to See: Collaborative Embodied Learning for Object Detection and Camera AdjustingLingdong Shen, Chunlei Huo, Nuo Xu, Chaowei Han et al.AAAI 2024 · 4 citations
Builds on8
- Situational Fusion of Visual Representation for Visual NavigationWilliam B. Shen, Danfei Xu, Yuke Zhu, Li Fei-Fei et al.ICCV 2019 · 70 citations
- Learning to Caption Images Through a Lifetime by Asking QuestionsTingke Shen, Amlan Kar, Sanja FidlerICCV 2019 · 33 citations
- When2com: Multi-Agent Perception via Communication Graph GroupingYen-Cheng Liu, Junjiao Tian, Nathaniel Glaser, Zsolt KiraCVPR 2020
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang et al.CVPR 2020
- SOON: Scenario Oriented Object Navigation With Graph-Based ExplorationFengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu et al.CVPR 2021
Related papers
- Just Ask: An Interactive Learning Framework for Vision and Language NavigationTa-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim et al.AAAI 2020 · 88 citations
- Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent DialoguesFrancesco Taioli, Edoardo Zorzi, Gianni Franchi, Alberto Castellini et al.ICCV 2025 · 1 citation
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied NavigationJuncheng Li, Xin Wang, Siliang Tang, Haizhou Shi et al.CVPR 2020
- A Framework for Learning to Request Rich and Contextually Useful Information from HumansKhanh X. Nguyen, Yonatan Bisk, Hal Daumé IIIICML 2022 · 21 citations
- CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy EnvironmentsXiulong Liu, Sudipta Paul, Moitreya Chatterjee, Anoop CherianAAAI 2024 · 16 citations
