4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models
Wanhua Li, Renping Zhou, Jiawei Zhou, Yingwei Song, Johannes Herter, Minghan Qin, Gao Huang, Hanspeter Pfister
Abstract
Figure 1. Visualization of the learned language features of our 4D LangSplat. We observe that 4D LangSplat effectively learns dynamic semantic features that change over time, such as the gradual diffusion of coffee shown in the first two rows, and the "chicken" toggling between open and closed states in the latter two rows. Additionally, our semantic field captures consistent features for semantics that remain unchanged over time, with the clear object boundaries in the visualization demonstrating the precision of our semantic field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fca9c9a2-372d-42ea-b91e-6ba1e68fcb3fCited by top-tier papers12
- LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPSWanhua Li, Yujie Zhao, Minghan Qin, Yang Liu et al.NeurIPS 2025 · 54 citations
- Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout DecodingYixiong Fang, Ziran Yang, Zhaorun Chen, Zhuokai Zhao et al.NeurIPS 2025 · 21 citations
- Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian SplattingKangjie Chen, Yingji Zhong, Zhihao Li, Jiaqi Lin et al.NeurIPS 2025 · 15 citations
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingKairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng et al.NeurIPS 2025 · 11 citations
- Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian SplattingYiren Lu, Yunlai Zhou, Yiran Qiao, Chaoda Song et al.NeurIPS 2025 · 9 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
Related papers
- LangField4D: Learning Identity-Adaptive and Spatio-Temporal Continuous 4D Language Fields for Dynamic ScenesYichao Xu, Qiaowei Miao, Jinsheng Quan, Wei Yang et al.CVPR 2026
- ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian SplattingYuming Meng, Dong Wu, Hongbin ZhaCVPR 2026
- LangSplat: 3D Language Gaussian SplattingMinghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang et al.CVPR 2024 · 164 citations
- Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene UnderstandingPedro Hermosilla, Christian Stippel, Leon SickCVPR 2025
- LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene UnderstandingHanyu Zhou, Gim Hee LeeICLR 2026 · 20 citations
