3D Vision-Language Gaussian Splatting
Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng, Anwesa Choudhuri, Terrence Chen, Chen Chen, Ziyan Wu
Abstract
Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches have naively embedded semantic representations into 3D reconstruction methods without striking a balance between visual and language modalities, which leads to unsatisfying semantic rasterization of translucent or reflective objects, as well as over-fitting on color modality. To alleviate these limitations, we propose a solution that adequately handles the distinct visual and semantic modalities, i.e., a 3D visionlanguage Gaussian splatting model for scene understanding, to put emphasis on the representation learning of language modality. We propose a novel cross-modal rasterizer, using modality fusion along with a smoothed semantic indicator for enhancing semantic rasterization. We also employ a camera-view blending technique to improve semantic consistency between existing and synthesized views, thereby effectively mitigating over-fitting. Extensive experiments demonstrate that our method achieves state-of-the-art performance in open-vocabulary semantic segmentation, surpassing existing methods by a significant margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene EncodingYue Li, Qi Ma, Runyi Yang, Mengjiao Ma et al.CVPR 2026 · 10 citations
- Splat Feature SolverButian Xiong, Rong Liu, Kenneth Xu, Meida Chen et al.ICLR 2026 · 7 citations
- OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene UnderstandingSheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng SunCVPR 2026 · 5 citations
- Universal Beta SplattingRong Liu, Zhongpai Gao, Benjamin Planche, Meida Chen et al.ICLR 2026 · 4 citations
- Splattalk: 3D VQA with Gaussian SplattingAnh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas et al.ICCV 2025 · 4 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- Tackling View-Dependent Semantics in 3D Language Gaussian SplattingJiazhong Cen, Xudong Zhou, Jiemin Fang, Changsong Wen et al.ICML 2025
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingXiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei et al.ICCV 2025 · 1 citation
- Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian SplattingYiren Lu, Yunlai Zhou, Yiran Qiao, Chaoda Song et al.NeurIPS 2025 · 9 citations
- VaF-LangSplat: Voxel-Aware Fusion Language Gaussian SplattingChangzhou Li, Xinyu Yang, Weiguo Yang, Xinyi LiACM MM 2025
- GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space HyperplaneYansong Qu, Shaohui Dai, Xinyang Li, Jianghang Lin et al.ACM MM 2024 · 20 citations
