Vision-Dialog Navigation by Exploring Cross-Modal Memory
Yi Zhu, Fengda Zhu, Zhaohuan Zhan, Bingqian Lin, Jianbin Jiao, Xiaojun Chang, Xiaodan Liang
Abstract
Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language navigation, vision-dialog navigation also requires to handle well with the language intentions of a series of questions about the temporal context from dialogue history and co-reasoning both dialogs and visual scenes. In this paper, we propose the Cross-modal Memory Network (CMN) for remembering and understanding the rich information relevant to historical navigation actions. Our CMN consists of two memory modules, the language memory module (L-mem) and the visual memory module (V-mem). Specifically, L-mem learns latent relationships between the current language interaction and a dialog history by employing a multi-head attention mechanism. V-mem learns to associate the current visual views and the cross-modal memory about the previous navigation actions. The crossmodal memory is generated via a vision-to-language attention and a language-to-vision attention. Benefiting from the collaborative learning of the L-mem and the V-mem, our CMN is able to explore the memory about the decision making of historical navigation actions which is for the current step. Experiments on the CVDN dataset show that our CMN outperforms the previous state-of-the-art model by a significant margin on both seen and unseen environments. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84edaaa3-8231-4a60-8c5b-893c2ff9d509Cited by top-tier papers10
- The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language NavigationYuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang et al.ICCV 2021 · 87 citations
- HOP: History-and-Order Aware Pretraining for Vision-and-Language NavigationYanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu et al.CVPR 2022 · 71 citations
- Self-Motivated Communication Agent for Real-World Vision-Dialog NavigationYi Zhu, Yue Weng, Fengda Zhu, Xiaodan Liang et al.ICCV 2021 · 41 citations
- Frequency-Enhanced Data Augmentation for Vision-and-Language NavigationKeji He, Chenyang Si, Zhihe Lu, Yan Huang et al.NeurIPS 2023 · 32 citations
- Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationYibo Cui, Liang Xie, Yakun Zhang, Meishan Zhang et al.ICCV 2023 · 31 citations
Builds on1
Related papers
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Cross-modal Map Learning for Vision and Language NavigationGeorgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan et al.CVPR 2022 · 2 citations
- NDH-Full: Learning and Evaluating Navigational Agents on Full-Length DialogueHyounghun Kim, Jialu Li, Mohit BansalEMNLP 2021 · 11 citations
- Learning Fine-Grained Alignment for Aerial Vision-Dialog NavigationYifei Su, Dong An, Kehan Chen, Weichen Yu et al.AAAI 2025 · 7 citations
- KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueXiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun et al.ACM MM 2020 · 37 citations
