Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
Akhil Perincherry, Jacob Krantz, Stefan Lee
摘要
Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we study if visual representations of sub-goals implied by the instructions can serve as navigational cues and lead to increased navigation performance. To synthesize these visual representations or "imaginations", we leverage a text-to-image diffusion model on landmark references contained in segmented instructions. These imaginations are provided to VLN agents as an added modality to act as landmark cues and an auxiliary loss is added to explicitly encourage relating these with their corresponding referring expressions. Our findings reveal an increase in success rate (SR) of ∼1 point and up to ∼0.5 points in success scaled by inverse path length (SPL) across agents. These results suggest that the proposed approach reinforces visual understanding compared to relying on language instructions alone. Code and data for our work can be found at https://www.akhilperincherry . com/VLN-Imagine-website/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Uncertainty-Aware Gaussian Map for Vision-Language NavigationJianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao 等ICLR 2026 · 被引用 3 次
- Harnessing Input-Adaptive Inference for Efficient VLNDongwoo Kang, Akhil Perincherry, Zachary Coalson, Aiden Gabriel 等ICCV 2025 · 被引用 1 次
- Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual NavigationNingnan Wang, Weihuang Chen, Liming Chen, Haoxuan Ji 等AAAI 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
相关 Paper
- Vision-Language Navigation With Self-Supervised Auxiliary Reasoning TasksFengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan LiangCVPR 2020
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 被引用 29 次
- Improving Vision-and-Language Navigation by Generating Future-View Image SemanticsJialu Li, Mohit BansalCVPR 2023
- SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language NavigationAbhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee 等NeurIPS 2021 · 被引用 88 次
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu 等NeurIPS 2020 · 被引用 167 次
