GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting
Yuning Peng, Haiping Wang, Yuan Liu, Chenglu Wen, Zhen Dong, Bisheng Yang
摘要
3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary queries for renderings on arbitrary viewpoints. The main challenge of distilling 2D features for 3D fields lies in the multiview inconsistency of extracted 2D features, which provides unstable supervision for the 3D feature field. GAGS addresses this challenge with two novel strategies. First, GAGS associates the prompt point density of SAM with the camera distances to scene objects, which significantly improves the multiview consistency of segmentation results. Second, GAGS further decodes a granularity factor to guide the distillation process and this granularity factor can be learned in a unsupervised manner to only select the multiview consistent 2D features in the distillation process. Experimental results on two datasets show that GAGS improves visual grounding accuracy by an average of 10.9% and semantic segmentation accuracy by an average of 7.0%, with an inference speed 2× faster than baseline methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian SplattingChangyue Shi, Minghao Chen, Yiping Mao, Chuxiao Yang 等CVPR 2026 · 被引用 8 次
- Splat Feature SolverButian Xiong, Rong Liu, Kenneth Xu, Meida Chen 等ICLR 2026 · 被引用 7 次
- RoboPearls: Editable Video Simulation for Robot ManipulationTang Tao, Likui Zhang, Youpeng Wen, Kaidong Zhang 等ICCV 2025 · 被引用 4 次
- ObjectGS: Object-Aware Scene Reconstruction and Scene Understanding via Gaussian SplattingRuijie Zhu, Mulin Yu, Linning Xu, Lihan Jiang 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
相关 Paper
- Rh-3DGS: Robust Open-Vocabulary Scene Understanding via Riemannian Huber Distillation and Manifold-Aware SamplingXinpeng Zhao, Jiang Jie, Fengyuan Zhang, Lixin Zhan 等ICML 2026
- Tackling View-Dependent Semantics in 3D Language Gaussian SplattingJiazhong Cen, Xudong Zhou, Jiemin Fang, Changsong Wen 等ICML 2025
- OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary UnderstandingYanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu 等NeurIPS 2024 · 被引用 191 次
- Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature FieldsShijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan 等CVPR 2024 · 被引用 145 次
- FHGS: Feature-Homogenized Gaussian SplattingQigeng Duan, Benyun Zhao, Mingqiao Han, Yijun Huang 等NeurIPS 2025 · 被引用 2 次
