Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation
Yiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng Wang
Abstract
Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of embodied agents in unseen environments. However, existing Vision-and-Language Navigation (VLN) agents lack a memory mechanism of this kind. To address this, we propose a novel architecture that equips agents with a reality-imagination hybrid memory system. This system enables agents to maintain and expand their memory through both imaginative mechanisms and navigation actions. Additionally, we design tailored pre-training tasks to develop the agent's imaginative capabilities. Our agent can imagine high-fidelity RGB images for future scenes, achieving state-of-the-art results in a Success rate weighted by Path Length (SPL).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03de6c0c-aacb-41fa-9819-0ccff26df570Cited by top-tier papers6
- Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language NavigationPingrui Zhang, Yifei Su, Pengyuan Wu, Dong An et al.CVPR 2026 · 19 citations
- Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual NavigationYiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng WangNeurIPS 2025 · 5 citations
- An Effective Levelling Paradigm for Unlabeled ScenariosFangming Cui, Zhou Yu, Di Yang, Yuqiang Ren et al.NeurIPS 2025
- Enhancing Target-unspecific Tasks through a Features MatrixFangming Cui, Yonggang Zhang, Xuan Wang, Xinmei Tian et al.ICML 2025
- FloVerse: Floor Plan-Guided Multi-Modal NavigationWeiqi Huang, Shuangyi Dong, Jiaxin Li, Yifei Guo et al.CVPR 2026
Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
Related papers
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language NavigationShuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong et al.ICLR 2026 · 124 citations
- Pathdreamer: A World Model for Indoor NavigationJing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge et al.ICCV 2021 · 128 citations
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li et al.CVPR 2026 · 7 citations
- Counterfactual Vision-and-Language Navigation: Unravelling the UnseenAmin Parvaneh, Ehsan Abbasnejad, Damien Teney, Qinfeng Shi et al.NeurIPS 2020 · 54 citations
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation PredictionFei Liu, Shichao Xie, Minghua Luo, Zedong Chu et al.CVPR 2026 · 16 citations
