From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang, Ayush Jain, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson
摘要
3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes-a six-order-ofmagnitude gap that severely limits performance. We introduce LIFT-GS, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This rendersupervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with 25.7% mAP on open-vocabulary instance segmentation (vs. 20.2% prior SOTA) and consistent 10-30% improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies finetuning datasets by 2×, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene EncodingYue Li, Qi Ma, Runyi Yang, Mengjiao Ma 等CVPR 2026 · 被引用 10 次
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance FieldsLisa Weijler, Sebastian Koch, Fabio Poiesi, Timo Ropinski 等NeurIPS 2025 · 被引用 5 次
- GenSplat: Bridging the Generalization Gap in 3DGS Language ComprehensionFang Liu, Yuhao Liu, Ke Xu, Gerhard Hancke 等CVPR 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang 等CVPR 2026
- Unifying 2D and 3D Vision-Language UnderstandingAyush Jain, Alexander Swerdlow, Yuzhou Wang, Sergio Arnaud 等ICML 2025
- GAGS: Granularity-Aware Feature Distillation for Language Gaussian SplattingYuning Peng, Haiping Wang, Yuan Liu, Chenglu Wen 等AAAI 2026 · 被引用 12 次
- Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation ModelsZhimin Chen, Longlong Jing, Yingwei Li, Bing LiNeurIPS 2023 · 被引用 57 次
- Weakly Supervised 3D Open-vocabulary SegmentationKunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu 等NeurIPS 2023 · 被引用 173 次
