MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
Shuo Wang, Yongcai Wang, Zhaoxin Fan, Yucheng Wang, Maiyue Chen, Kaihui Wang, Zhizhong Su, Wanting Li, Xudong Cai, Yeying Jin, Deying Li
Abstract
Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results with monocular input, yet they still lag behind methods using panoramic RGB-D information. We present MonoDream, a lightweight VLA framework that enables monocular agents to learn a Unified Navigation Representation (UNR). This shared feature representation jointly aligns navigation-relevant visual semantics (e.g., global layout, depth, and future cues) and language-grounded action intent, enabling more reliable action prediction. MonoDream further introduces Latent Panoramic Dreaming (LPD) tasks to supervise the UNR, which train the model to predict latent features of panoramic RGB and depth observations at both current and future steps based on only monocular input. Experiments on multiple VLN benchmarks show that MonoDream consistently improves monocular navigation performance and significantly narrows the gap with panoramic-based agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07ede432-c01e-4733-b7da-dad2b2e07a27Cited by top-tier papers4
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language NavigationShuo Wang, Yongcai Wang, Wanting Li, Xudong Cai et al.NeurIPS 2025 · 28 citations
- Progress-Think: Semantic Progress Reasoning for Vision-Language NavigationShuo Wang, Yucheng Wang, Guoxin Lian, Yongcai Wang et al.CVPR 2026 · 10 citations
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and NavigationZihan Wang, Seungjun Lee, Guangzhao Dai, Gim Hee LeeCVPR 2026 · 9 citations
- GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language NavigationJiahao Yang, Zihan Wang, Xiangyang Li, Xing Zhu et al.CVPR 2026 · 1 citation
Builds on19
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen et al.ICLR 2024 · 309 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
Related papers
- MapDream: Task-Driven Map Learning for Vision-Language NavigationGuoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang et al.ICML 2026
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 25 citations
- monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language NavigationRenjie Lu, Yu Zhou, Hao Cheng, Jingke Meng et al.ICCV 2025 · 1 citation
- Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.CVPR 2024 · 13 citations
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li et al.CVPR 2026 · 7 citations
