From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang, Ayush Jain, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson
Abstract
3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes-a six-order-ofmagnitude gap that severely limits performance. We introduce LIFT-GS, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This rendersupervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with 25.7% mAP on open-vocabulary instance segmentation (vs. 20.2% prior SOTA) and consistent 10-30% improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies finetuning datasets by 2×, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48af6677-4757-44a6-930c-e86c8e9cfe56Cited by top-tier papers3
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene EncodingYue Li, Qi Ma, Runyi Yang, Mengjiao Ma et al.CVPR 2026 · 10 citations
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance FieldsLisa Weijler, Sebastian Koch, Fabio Poiesi, Timo Ropinski et al.NeurIPS 2025 · 5 citations
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke et al.CVPR 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang et al.CVPR 2026
- Unifying 2D and 3D Vision-Language UnderstandingAyush Jain, Alexander Swerdlow, Yuzhou Wang, Sergio Arnaud et al.ICML 2025
- GAGS: Granularity-Aware Feature Distillation for Language Gaussian SplattingYuning Peng, Haiping Wang, Yuan Liu, Chenglu Wen et al.AAAI 2026 · 12 citations
- Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation ModelsZhimin Chen, Longlong Jing, Yingwei Li, Bing LiNeurIPS 2023 · 57 citations
- Weakly Supervised 3D Open-vocabulary SegmentationKunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu et al.NeurIPS 2023 · 173 citations
