DiaLoc: An Iterative Approach to Embodied Dialog Localization
Chao Zhang, Mohan Li, Ignas Budvytis, Stephan Liwicki
Abstract
Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing dialogbased localization approaches assume the availability of entire dialog prior to Iocalizaiton, which is impractical for deployed dialog-based localization. In this paper, we propose DiaLoc, a new dialog-based localization framework which aligns with a real human operator behavior. Specifically, we produce an iterative refinement of location predictions which can visualize current pose believes after each dialog turn. DiaLoc effectively utilizes the multimodal data for multi-shot localization, where a fusion encoder fuses vision and dialog information iteratively. We achieve state-of-the-art results on embodied dialog-based localization task, in single-shot (+7.08% in Acc5@valUnseen) and multishot settings (+10.85% in Acc5@valUnseen). DiaLoc narrows the gap between simulation and real-world applications, opening doors for future research on collaborative localization and navigation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 763ae3c0-a53f-4d9e-87d6-cd5096f62fe7Cited by top-tier papers3
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
- DialNav: Multi-Turn Dialog Navigation with a Remote GuideLeekyeung Han, Hyunji Min, Gyeom Hwangbo, Jonghyun Choi et al.ICCV 2025 · 1 citation
- Towards Precise Embodied Dialogue Localization via Causality Guided DiffusionHaoyu Wang, Le Wang, Sanping Zhou, Jingyi Tian et al.CVPR 2025
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- Where Are You? Localization from Embodied DialogMeera Hahn, Jacob Krantz, Dhruv Batra, Devi Parikh et al.EMNLP 2020 · 22 citations
- Conversational Localization: Indoor Human Localization through Intelligent ConversationSmitha Sheshadri, Kotaro HaraUbiComp 2024 · 5 citations
- DialogueVPR: Towards Conversational Visual Place RecognitionYukun Song, Changwei Wang, Xingtian Pei, Shibiao Xu et al.CVPR 2026 · 1 citation
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu et al.AAAI 2024 · 122 citations
- Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionGeorgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage et al.EMNLP 2023 · 1 citation
