Online Language Splatting
Saimouli Katragadda, Cho-Ying Wu, Yuliang Guo, Xinyu Huang, Guoquan Huang, Liu Ren
Abstract
To enable AI agents to interact seamlessly with both humans and 3D environments, they must not only perceive the 3D world accurately but also align human language with 3D spatial representations. While prior work has made significant progress by integrating language features into geometrically detailed 3D scene representations using 3D Gaussian Splatting (GS), these approaches rely on computationally intensive offline preprocessing of language features for each input image, limiting adaptability to new environments. In this work, we introduce Online Language Splatting, the first framework to achieve online, near real-time, open-vocabulary language mapping within a 3DGS-SLAM system without requiring pre-generated language features. The key challenge lies in efficiently fusing high-dimensional language features into 3D representations while balancing the computation speed, memory usage, rendering quality and open-vocabulary capability. To this end, we innovatively design: (1) a high-resolution CLIP embedding module capable of generating detailed language feature maps in 18ms per frame, (2) a two-stage online auto-encoder that compresses 768-dimensional CLIP features to 15 dimensions while preserving open-vocabulary capabilities, and (3) a color-language disentangled optimization approach to improve rendering quality. Experimental results show that our online method not only surpasses the state-of-the-art offline methods in accuracy but also achieves more than 40x efficiency boost, demonstrating the potential for dynamic and interactive AI applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75c32f0f-94f3-42cb-9996-19410cfbdb60Cited by top-tier papers3
- EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene UnderstandingSeungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee LeeCVPR 2026 · 2 citations
- No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D ConsistencyCho-Ying Wu, Zixun Huang, Xinyu Huang, Liu RenCVPR 2026
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke et al.CVPR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPSWanhua Li, Yujie Zhao, Minghan Qin, Yang Liu et al.NeurIPS 2025 · 54 citations
- VaF-LangSplat: Voxel-Aware Fusion Language Gaussian SplattingChangzhou Li, Xinyu Yang, Weiguo Yang, Xinyi LiACM MM 2025
- Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding RegistrationKim Jun-Seong, GeonU Kim, Kim Yu-Ji, Yu-Chiang Frank Wang et al.CVPR 2025
- FastLGS: Speeding Up Language Embedded Gaussians with Feature Grid MappingYuzhou Ji, He Zhu, Junshu Tang, Wuyi Liu et al.AAAI 2025 · 29 citations
- LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian SplattingXulun Ye, Qin Zhang, Kun ZhouCVPR 2026
