3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
Jianzhe Gao, Rui Liu, Wenguan Wang
摘要
Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough scene understanding. While existing works equip agents with various scene representations to enhance spatial awareness, they often neglect the complex 3D geometry and rich semantics in VLN scenarios, limiting the ability to generalize across diverse and unseen environments. To address these challenges, this work proposes a 3D Gaussian Map that represents the environment as a set of differentiable 3D Gaussians and accordingly develops a navigation strategy for VLN. Specifically, Egocentric Scene Map is constructed online by initializing 3D Gaussians from sparse pseudo-lidar point clouds, providing informative geometric priors for scene understanding. Each Gaussian primitive is further enriched through Open-Set Semantic Grouping operation, which groups 3D Gaussians based on their membership in object instances or stuff categories within the open world, resulting in a unified 3D Gaussian Map. Building on this map, Multi-Level Action Prediction strategy, which combines spatial-semantic cues at multiple granularities, is designed to assist agents in decision-making. Extensive experiments conducted on three public benchmarks (i.e., R2R, R4R, and REVERIE) validate the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Underwater Visual SLAM with Depth Uncertainty and Medium ModelingRui Liu, Sheng Fan, Wenguan Wang, Yi YangICCV 2025 · 被引用 6 次
- Uncertainty-Aware Gaussian Map for Vision-Language NavigationJianzhe Gao, Rui Liu, Yuxuan Xu, Tongtong Cao 等ICLR 2026 · 被引用 3 次
- Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and ReasoningSixian Zhang, Yiyao Wang, Xinhang Song, Keming Zhang 等CVPR 2026 · 被引用 3 次
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language NavigationXichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang 等AAAI 2026 · 被引用 3 次
- Gaussian-Based World Model: Gaussian Priors for Voxel-Based Occupancy Prediction and Future Motion PredictionTuo Feng, Wenguan Wang, Yi YangICCV 2025 · 被引用 1 次
它引用的顶会 Paper56
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 被引用 551 次
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 被引用 522 次
相关 Paper
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 被引用 25 次
- Bird's-Eye-View Scene Graph for Vision-Language NavigationRui Liu, Xiaohan Wang, Wenguan Wang, Yi YangICCV 2023 · 被引用 100 次
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等ICCV 2023 · 被引用 136 次
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language NavigationKailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu 等CVPR 2026 · 被引用 10 次
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language NavigationZihan Wang, Seungjun Lee, Gim Hee LeeNeurIPS 2025 · 被引用 36 次
