InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception
Haijie Li, Yanmin Wu, Jiarui Meng, Qiankun Gao, Zhiyao Zhang, Ronggang Wang, Jian Zhang
Abstract
3D scene understanding is vital for applications in autonomous driving, robotics, and augmented reality. However, scene understanding based on 3D Gaussian Splatting faces three key challenges: (i) an imbalance between appearance and semantics, (ii) inconsistencies in object boundaries, and (iii) difficulties with top-down instance segmentation. To address these challenges, we propose InstanceGaussian, a method that jointly learns appearance and semantic features while adaptively aggregating instances. Our contributions are as follows: (i) a new Semantic-Scaffold-GS representation to improve feature representation and boundary delineation, (ii) a progressive training strategy for enhanced stability and segmentation, and (iii) a category-agnostic, bottom-up instance aggregation approach for better segmentation. Experimental results demonstrate that our approach achieves state-of-the-art performance in category-agnostic, openvocabulary 3D point-level segmentation, validating the effectiveness of our proposed method. Project page: https://lhj-git.github.io/InstanceGaussian/ * Corresponding author This work is supported by Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology(Grant No. 2024B1212010006) * The term "semantic" is used here to distinguish it from "appearance". In subsequent sections and code implementations, this concept will be explicitly referred to as "instance features".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Towards Physically Executable 3D Gaussian for Embodied NavigationBingchen Miao, Rong Wei, Zhiqi Ge, Xiaoquan sun et al.ICLR 2026 · 22 citations
- COS3D: Collaborative Open-Vocabulary 3D SegmentationRunsong Zhu, Ka-Hei Hui, Zhengzhe Liu, Qianyi Wu et al.NeurIPS 2025 · 12 citations
- ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian SplattingJiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long et al.CVPR 2026 · 7 citations
- EA3D: Online Open-World 3D Object Extraction from Streaming VideosXiaoyu Zhou, Jingqi Wang, Yuang Jia, Yongtao Wang et al.NeurIPS 2025 · 5 citations
- Splattalk: 3D VQA with Gaussian SplattingAnh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas et al.ICCV 2025 · 4 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary UnderstandingYanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu et al.NeurIPS 2024 · 191 citations
- Tackling View-Dependent Semantics in 3D Language Gaussian SplattingJiazhong Cen, Xudong Zhou, Jiemin Fang, Changsong Wen et al.ICML 2025
- Votesplat: Hough Voting Gaussian Splatting for 3D Scene UnderstandingMinchao Jiang, Shunyu Jia, Jiaming Gu, Xiaoyuan Lu et al.ICCV 2025 · 1 citation
- VaF-LangSplat: Voxel-Aware Fusion Language Gaussian SplattingChangzhou Li, Xinyu Yang, Weiguo Yang, Xinyi LiACM MM 2025
- ObjectGS: Object-Aware Scene Reconstruction and Scene Understanding via Gaussian SplattingRuijie Zhu, Mulin Yu, Linning Xu, Lihan Jiang et al.ICCV 2025 · 1 citation
