LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
Fangfu Liu, Hao Li, Jiawei Chi, Hanyang Wang, Ming-Hsuan Yang, Fudong Wang, Yueqi Duan
Abstract
Recovering 3D structures with open-vocabulary scene understanding from 2D images is a fundamental but daunting task. Recent developments have achieved this by performing per-scene optimization with embedded language information. However, they heavily rely on the calibrated denseview reconstruction paradigm, thereby suffering from severe rendering artifacts and implausible semantic synthesis when limited views are available. In this paper, we introduce a novel generative framework, coined LangScene-X, to unify and generate 3D consistent multi-modality information for reconstruction and understanding. Powered by the generative capability of creating more consistent novel observations, we can build generalizable 3D language-embedded scenes from only sparse views. Specifically, we first train a TriMap video diffusion model that can generate appearance (RGBs), geometry (normals), and semantics (segmentation maps) from sparse inputs through progressive knowledge integration. Furthermore, we propose a Language Quantized Compressor (LQC), trained on largescale image datasets, to efficiently encode language embeddings, enabling cross-scene generalization without perscene retraining. Finally, we reconstruct the language surface fields by aligning language information onto the surface of 3D scenes, enabling open-ended language queries. Extensive experiments on real-world data demonstrate the superiority of our LangScene-X over state-of-the-art methods in terms of quality and generalizability. Project Page: https://liuff19.github.io/LangScene-X/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Part-X-MLLM: Part-aware 3D Multimodal Large Language ModelChunshi Wang, Junliang Ye, Yunhan Yang, YANG LI et al.ICLR 2026 · 6 citations
- CFG-Ctrl: Control-Based Classifier-Free Diffusion GuidanceHanyang Wang, Yiyang Liu, Jiawei Chi, Fangfu Liu et al.CVPR 2026 · 5 citations
- NG-GS: NeRF-guided 3D Gaussian Splatting SegmentationYi He, Tao Wang, Yi Jin, Congyan Lang et al.CVPR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
Related papers
- Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding RegistrationKim Jun-Seong, GeonU Kim, Kim Yu-Ji, Yu-Chiang Frank Wang et al.CVPR 2025
- Taking Language Embedded 3D Gaussian Splatting into the WildYuze Wang, Junyi Wang, Yue QiIEEE VR 2026
- Rethinking Open-Vocabulary Segmentation of Radiance Fields in 3D SpaceHyunjee Lee, Youngsik Yun, Jeongmin Bae, Seoha Kim et al.AAAI 2025 · 11 citations
- ReLaGS: Relational Language Gaussian SplattingYaxu Xie, Abdalla Arafa, Alireza Javanmardi, Christen Millerdurai et al.CVPR 2026 · 7 citations
- Language Embedded 3D Gaussians for Open-Vocabulary Scene UnderstandingJin-Chuan Shi, Miao Wang, Hao-Bin Duan, Shao-Hua GuanCVPR 2024 · 57 citations
