Rethinking Open-Vocabulary Segmentation of Radiance Fields in 3D Space
Hyunjee Lee, Youngsik Yun, Jeongmin Bae, Seoha Kim, Youngjung Uh
Abstract
Understanding the 3D semantics of a scene is a fundamental problem for various scenarios such as embodied agents. While NeRFs and 3DGS excel at novel-view synthesis, previous methods for understanding their semantics have been limited to incomplete 3D understanding: their segmentation results are rendered as 2D masks that do not represent the entire 3D space. To address this limitation, we redefine the problem to segment the 3D volume and propose the following methods for better 3D understanding. We directly supervise the 3D points to train the language embedding field, unlike previous methods that anchor supervision at 2D pixels. We transfer the learned language field to 3DGS, achieving the first real-time rendering speed without sacrificing training time or accuracy. Lastly, we introduce a 3D querying and evaluation protocol for assessing the reconstructed geometry and semantics together. Code, checkpoints, and annotations are available at the project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1df0abfd-97ad-4eac-8fef-4d30b4707ad6Cited by top-tier papers7
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance TracingHongyu Shen, Junfeng Ni, Yixin Chen, Weishuo Li et al.ICCV 2025 · 5 citations
- Splat and Replace: 3D Reconstruction with Repetitive ElementsNicolás Violante, Andreas Meuleman, Alban Gauthier, Frédo Durand et al.SIGGRAPH 2025 · 4 citations
- Tackling View-Dependent Semantics in 3D Language Gaussian SplattingJiazhong Cen, Xudong Zhou, Jiemin Fang, Changsong Wen et al.ICML 2025
- Rh-3DGS: Robust Open-Vocabulary Scene Understanding via Riemannian Huber Distillation and Manifold-Aware SamplingXinpeng Zhao, Jiang Jie, Fengyuan Zhang, Lixin Zhan et al.ICML 2026
- LangField4D: Learning Identity-Adaptive and Spatio-Temporal Continuous 4D Language Fields for Dynamic ScenesYichao Xu, Qiaowei Miao, Jinsheng Quan, Wei Yang et al.CVPR 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint GraphsShaohui Dai, Yansong Qu, Zheyan Li, Xinyang Li et al.ACM MM 2025 · 3 citations
- EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene UnderstandingSeungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee LeeCVPR 2026 · 2 citations
- LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian SplattingXulun Ye, Qin Zhang, Kun ZhouCVPR 2026
- ObjectGS: Object-Aware Scene Reconstruction and Scene Understanding via Gaussian SplattingRuijie Zhu, Mulin Yu, Linning Xu, Lihan Jiang et al.ICCV 2025 · 1 citation
- GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic FieldsYunsong Wang, Hanlin Chen, Gim Hee LeeCVPR 2024 · 2 citations
