EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
Seungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee Lee
Abstract
Understanding a 3D scene immediately with its exploration is essential for embodied tasks, where an agent must construct and comprehend the 3D scene in an online and nearly real-time manner. In this study, we propose EmbodiedSplat, an online feed-forward 3DGS for open-vocabulary scene understanding that enables simultaneous online 3D reconstruction and 3D semantic understanding from the streaming images. Unlike existing open-vocabulary 3DGS methods which are typically restricted to either offline or per-scene optimization setting, our objectives are two-fold: 1) Reconstructs the semantic-embedded 3DGS of the entire scene from over 300 streaming images in an online manner. 2) Highly generalizable to novel scenes with feed-forward design and supports nearly real-time 3D semantic reconstruction when combined with real-time 2D models. To achieve these objectives, we propose an Online Sparse Coefficients Field with a CLIP Global Codebook where it binds the 2D CLIP embeddings to each 3D Gaussian while minimizing memory consumption and preserving the full semantic generalizability of CLIP. Furthermore, we generate 3D geometric-aware CLIP features by aggregating the partial point cloud of 3DGS through 3D U-Net to compensate the 3D geometric prior to 2D-oriented language embeddings. Extensive experiments on diverse indoor datasets, including ScanNet, ScanNet++, and Replica, demonstrate both the effectiveness and efficiency of our method. Check out our project page in https://0nandon.github.io/EmbodiedSplat/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17e52430-e8d2-4707-a712-ca3f61f178f0Cited by top-tier papers1
Ask how each one uses itBuilds on48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
Related papers
- OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian SplattingHongjia Zhai, Qi Zhang, Xiaokun Pan, Xiyu Zhang et al.CVPR 2026 · 3 citations
- Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding RegistrationKim Jun-Seong, GeonU Kim, Kim Yu-Ji, Yu-Chiang Frank Wang et al.CVPR 2025
- SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language PretrainingYue Li, Qi Ma, Runyi Yang, Huapeng Li et al.ICCV 2025 · 5 citations
- Online Language SplattingSaimouli Katragadda, Cho-Ying Wu, Yuliang Guo, Xinyu Huang et al.ICCV 2025 · 2 citations
- S2GS: Streaming Semantic Gaussian Splatting for Online Scene Understanding and ReconstructionRenhe Zhang, Yuyang Tan, Jingyu Gong, Zhizhong Zhang et al.ICML 2026
