Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation
David Shavin, Sagie Benaim
摘要
Vision Foundation Models (VFMs) have achieved remarkable success when applied to various downstream 2D tasks. Despite their effectiveness, they often exhibit a critical lack of 3D awareness. To this end, we introduce Splat and Distill, a framework that instills robust 3D awareness into 2D VFMs by augmenting the teacher model with a fast, feed-forward 3D reconstruction pipeline. Given 2D features produced by a teacher model, our method first lifts these features into an explicit 3D Gaussian representation, in a feedforward manner. These 3D features are then "splatted" onto novel viewpoints, producing a set of novel 2D feature maps used to supervise the student model, "distilling" geometrically grounded knowledge. By replacing slow per-scene optimization of prior work with our feed-forward lifting approach, our framework avoids feature-averaging artifacts, creating a dynamic learning process where the teacher’s consistency improves alongside that of the student. We conduct a comprehensive evaluation on a suite of downstream tasks, including monocular depth estimation, surface normal estimation, multi-view correspondence, and semantic segmentation. Our method significantly outperforms prior works, not only achieving substantial gains in 3D awareness but also enhancing the underlying semantic richness of 2D features. Our project page and code are available here
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature FieldsShijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan 等CVPR 2024 · 被引用 145 次
- Feat2GS: Probing Visual Foundation Models with Gaussian SplattingYue Chen, Xingyu Chen, Anpei Chen, Gerard Pons-Moll 等CVPR 2025
- LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting ScenesJuliette Marrie, Romain Menegaux, Michael Arbel, Diane Larlus 等ICCV 2025 · 被引用 3 次
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsYifan Liu, Keyu Fan, Weihao Yu, Chenxin Li 等CVPR 2025
- FHGS: Feature-Homogenized Gaussian SplattingQigeng Duan, Benyun Zhao, Mingqiao Han, Yijun Huang 等NeurIPS 2025 · 被引用 2 次
