Multimodal LLM Guided Exploration and Active Mapping Using Fisher Information
Wen Jiang, Boshu Lei, Katrina Ashton, Kostas Daniilidis
Abstract
We present an active mapping system which plans for both long-horizon exploration goals and short-term actions using a 3D Gaussian Splatting (3DGS) representation. Existing methods either do not take advantage of recent developments in multimodal Large Language Models (LLM) or do not consider challenges in localization uncertainty, which is critical in embodied agents. We propose employing multimodal LLMs for long-horizon planning in conjunction with detailed motion planning using our information-based objective. By leveraging high-quality view synthesis from our 3DGS representation, our method employs a multimodal LLM as a zero-shot planner for long-horizon exploration goals from the semantic perspective. We also introduce an uncertainty-aware path proposal and selection algorithm that balances the dual objectives of maximizing the information gain for the environment while minimizing the cost of localization errors. Experiments conducted on the Gibson and Habitat-Matterport 3D datasets demonstrate state-of-the-art results of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b8a59b9-b93e-4d3e-a70d-edb84315e121Cited by top-tier papers4
- ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based ModelBoshu Lei, Wen Jiang, Kostas DaniilidisCVPR 2026 · 5 citations
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language GuidanceTianling Xu, Shengzhe Gan, Leslie Gu, Yuelei Li et al.CVPR 2026 · 5 citations
- Understanding while Exploring: Semantics-driven Active MappingLiyan Chen, Huangying Zhan, Hairong Yin, Yi Xu et al.NeurIPS 2025 · 5 citations
- Zero-shot Active Mapping via Fused 360-BEV Representations and Vision–Language ModelsYuanze Wang, Dianxi Shi, Yuetian Wang, Shiming Song et al.ICML 2026
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
Related papers
- Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and ReasoningSixian Zhang, Yiyao Wang, Xinhang Song, Keming Zhang et al.CVPR 2026 · 3 citations
- Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision GroundingJiaxin Shi, Mingyue Xiang, Hao Sun, Yixuan Huang et al.CVPR 2025
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke et al.CVPR 2026
- BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object NavigationZibo Zhou, Yue Hu, Lingkai Zhang, Zonglin Li et al.NeurIPS 2025 · 31 citations
- ActiveGAMER: Active GAussian Mapping through Efficient RenderingLiyan Chen, Huangying Zhan, Kevin Chen, Xiangyu Xu et al.CVPR 2025
