Identity-Aware Language Gaussian Splatting for Open-Vocabulary 3D Semantic Segmentation
SungMin Jang, Wonjun Kim
Abstract
Open-vocabulary 3D semantic segmentation has been actively studied by incorporating language features into 3D scene representations. Even though many methods have shown the notable improvement in this task, they still have difficulties to make language embeddings be consistent across different views. This inconsistency highly results in mis-labeling where different language embeddings are assigned to the same part of an object. To address this issue, we propose a simple yet powerful method that aligns language embeddings via the identity information. The key idea is to locate language embeddings for the same identity closely in the latent space while putting them apart otherwise. This approach allows the same object to have identical language embeddings in novel views with accurate semantic masks, which are well aligned with the input text. Furthermore, we propose a progressive mask expanding scheme that enables more accurate extraction of semantic mask boundaries. This scheme is very effective in preserving the boundary shape of the target region by allowing the model to consider the local relationship between segments. Experimental results on benchmark datasets demonstrate that our method delivers state-of-the-art performance in open-vocabulary 3D semantic segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0b5014f-bacf-49f1-a45f-7de1c75a9f1eCited by top-tier papers5
- OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene UnderstandingSheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng SunCVPR 2026 · 5 citations
- Rh-3DGS: Robust Open-Vocabulary Scene Understanding via Riemannian Huber Distillation and Manifold-Aware SamplingXinpeng Zhao, Jiang Jie, Fengyuan Zhang, Lixin Zhan et al.ICML 2026
- LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian SplattingXulun Ye, Qin Zhang, Kun ZhouCVPR 2026
- BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object ExtractionAlessio Mazzucchelli, Maria Naranjo-Almeida, Jorge Bustos-Sanchez, Mariella Dimiccoli et al.CVPR 2026
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke et al.CVPR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
Related papers
- XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic SegmentationZiyi Wang, Yanbo Wang, Xumin Yu, Jie Zhou et al.NeurIPS 2024 · 7 citations
- Open-Vocabulary 3D Semantic Segmentation with Foundation ModelsLi Jiang, Shaoshuai Shi, Bernt SchieleCVPR 2024
- Aligning Bag of Regions for Open-Vocabulary Object DetectionSize Wu, Wenwei Zhang, Sheng Jin, Wentao Liu et al.CVPR 2023
- Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationYongkang Li, Tianheng Cheng, Bin Feng, Wenyu Liu et al.CVPR 2025
- Masked Point-Entity Contrast for Open-Vocabulary 3D Scene UnderstandingYan Wang, Baoxiong Jia, Ziyu Zhu, Siyuan HuangCVPR 2025
