Lune

ICML2026Top-tier venue

Zero-shot Active Mapping via Fused 360-BEV Representations and Vision–Language Models

Yuanze Wang, Dianxi Shi, Yuetian Wang, Shiming Song, Haikuo Peng, Chunping Qiu, Mengzhu Wang

2026Year

Abstract

Active mapping enables embodied agents to understand and interact in previously unseen environments. However, most methods struggle to achieve zero-shot generalization to large-scale scenes and lack support for language instructions. We propose a VLM-based active mapping method that achieves zero-shot mapping while facilitating language-driven human–agent interaction. First, we introduce a 360-BEV representation that integrates omnidirectional semantics with BEV-aligned geometric structure to enhance scene understanding. Second, we develop a candidate waypoint generation strategy that allows the VLM-driven agent to select informative 2D waypoints in image space and back-project them into executable metric actions in 3D space, enabling the VLM to plan in its strongest modality. Third, we design a VLM-based depth-first exploration agent that decomposes the scenes into explorable regions, selects informative waypoints within each region, and organizes them into a topological tree. The agent follows the depth-first exploration policy to achieve thorough coverage of large-scale scenes. Without task-specific training, our method outperforms the strongest baseline, improving coverage and AUC by approximately 13.25% and 14.00%, respectively, while enabling language-conditioned interaction.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 29b08e1e-9de8-44c0-aef3-cb28ef2a756d

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines