NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
Xuan Yao, Junyu Gao, Changsheng Xu
摘要
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing to novel environments and adapting to ongoing changes during navigation. Inspired by human cognition, we present NavMorph, a self-evolving world model framework that enhances environmental understanding and decision-making in VLN-CE tasks. NavMorph employs compact latent representations to model environmental dynamics, equipping agents with foresight for adaptive planning and policy refinement. By integrating a novel Contextual Evolution Memory, NavMorph leverages scene-contextual information to support effective navigation while maintaining online adaptability. Extensive experiments demonstrate that our method achieves notable performance improvements on popular VLN-CE benchmarks. Code is available at https://github.com/Feliciaxyao/NavMorph.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language NavigationShuang Zeng, Dekang Qi, Xinyuan Chang, Feng Xiong 等ICLR 2026 · 被引用 124 次
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation PredictionFei Liu, Shichao Xie, Minghua Luo, Zedong Chu 等CVPR 2026 · 被引用 16 次
- MonoDream: Monocular Vision-Language Navigation with Panoramic DreamingShuo Wang, Yongcai Wang, Zhaoxin Fan, Yucheng Wang 等AAAI 2026 · 被引用 11 次
- Progress-Think: Semantic Progress Reasoning for Vision-Language NavigationShuo Wang, Yucheng Wang, Guoxin Lian, Yongcai Wang 等CVPR 2026 · 被引用 10 次
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open WorldMingming Yu, Fei Zhu, Wenzhuo Liu, Yirong Yang 等NeurIPS 2025 · 被引用 10 次
它引用的顶会 Paper36
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang 等NeurIPS 2023 · 被引用 453 次
- Model-Based Imitation Learning for Urban DrivingAnthony Hu, Gianluca Corrado, Nicolas Griffiths, Zachary Murez 等NeurIPS 2022 · 被引用 241 次
相关 Paper
- SC-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous EnvironmentsXuan Yao, Yuze Zhu, JUNYU GAO, Zongmeng Wang 等ICML 2026
- History to Future: Evolving Agent with Experience and Thought for Zero-shot Vision-and-Language NavigationGuangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie 等CVPR 2026
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li 等CVPR 2026 · 被引用 7 次
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language NavigationZihan Wang, Seungjun Lee, Gim Hee LeeNeurIPS 2025 · 被引用 36 次
- Narrowing the Gap between Vision and Action in NavigationYue Zhang, Parisa KordjamshidiACM MM 2024 · 被引用 2 次
