Zero-shot Active Mapping via Fused 360-BEV Representations and Vision–Language Models
Yuanze Wang, Dianxi Shi, Yuetian Wang, Shiming Song, Haikuo Peng, Chunping Qiu, Mengzhu Wang
Abstract
Active mapping enables embodied agents to understand and interact in previously unseen environments. However, most methods struggle to achieve zero-shot generalization to large-scale scenes and lack support for language instructions. We propose a VLM-based active mapping method that achieves zero-shot mapping while facilitating language-driven human–agent interaction. First, we introduce a 360-BEV representation that integrates omnidirectional semantics with BEV-aligned geometric structure to enhance scene understanding. Second, we develop a candidate waypoint generation strategy that allows the VLM-driven agent to select informative 2D waypoints in image space and back-project them into executable metric actions in 3D space, enabling the VLM to plan in its strongest modality. Third, we design a VLM-based depth-first exploration agent that decomposes the scenes into explorable regions, selects informative waypoints within each region, and organizes them into a topological tree. The agent follows the depth-first exploration policy to achieve thorough coverage of large-scale scenes. Without task-specific training, our method outperforms the strongest baseline, improving coverage and AUC by approximately 13.25% and 14.00%, respectively, while enabling language-conditioned interaction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29b08e1e-9de8-44c0-aef3-cb28ef2a756dBuilds on12
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu et al.CVPR 2024 · 53 citations
- Active Neural MappingZike Yan, Haoxiang Yang, Hongbin ZhaICCV 2023 · 37 citations
Related papers
- Multimodal LLM Guided Exploration and Active Mapping Using Fisher InformationWen Jiang, Boshu Lei, Katrina Ashton, Kostas DaniilidisICCV 2025 · 4 citations
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language MappingDanyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang et al.ACM MM 2025 · 2 citations
- Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationJiaqi Chen, Bingqian Lin, Xinmin Liu, Lin Ma et al.AAAI 2025 · 61 citations
- Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile ManipulationTzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su, Chung-Yi Lin et al.AAAI 2026
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language NavigationKailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu et al.CVPR 2026 · 10 citations
