Foundry: Distilling 3D Foundation Models for the Edge
Guillaume Letellier, Siddharth Srivastava, Frédéric Jurie, Gaurav Sharma
Abstract
Foundation models pre-trained with self-supervised learning (SSL) on large-scale datasets have become powerful general-purpose feature extractors. However, their immense size and computational cost make them prohibitive for deployment on edge devices such as robots and AR/VR headsets. Existing compression techniques like standard knowledge distillation create efficient 'specialist' models but sacrifice the crucial, downstream-agnostic generality that makes foundation models so valuable.
In this paper, we introduce Foundation Model Distillation (FMD), a new paradigm for compressing large SSL models into compact, efficient, and faithful proxies that retain their general-purpose representational power. We present Foundry, the first implementation of FMD for 3D point clouds. Our approach, Foundry, trains a student to learn a compressed set of SuperTokens that reconstruct the teacher's token-level representations, capturing a compact basis of its latent space. A single distilled model maintains strong transferability across diverse downstream tasks-classification, part segmentation, and fewshot scenarios-approaching full foundation-model performance while using significantly fewer tokens and FLOPs, making such models more practical for deployment on resource-constrained hardware. Project page: https: //guigui14460.github.io/foundry.github. io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b25309ab-0a90-4daa-b8fc-08b741b8f711Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Efficient 3D Semantic Segmentation with Superpoint TransformerDamien Robert, Hugo Raguet, Loïc LandrieuICCV 2023 · 131 citations
Related papers
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan et al.CVPR 2026 · 1 citation
- SAM-Guided Masked Token Prediction for 3D Scene UnderstandingZhimin Chen, Liang Yang, Yingwei Li, Longlong Jing et al.NeurIPS 2024 · 12 citations
- From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMsAng Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang et al.ICML 2025
- Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy HessiansIshan Amin, Sanjeev Raja, Aditi S. KrishnapriyanICLR 2025 · 5 citations
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta et al.ICML 2024 · 15 citations
