Generic-to-Specific Distillation of Masked Autoencoders
Wei Huang, Zhiliang Peng, Li Dong, Furu Wei, Jianbin Jiao, Qixiang Ye
Abstract
Large vision Transformers (ViTs) driven by selfsupervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, benefit little from those pretraining mechanisms. Knowledge distillation defines a paradigm to transfer representations from large (teacher) models to small (student) ones. However, the conventional single-stage distillation easily gets stuck on task-specific transfer, failing to retain the task-agnostic knowledge crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to tap the potential of small ViT models under the supervision of large models pre-trained by masked autoencoders. In generic distillation, decoder of the small model is encouraged to align feature predictions with hidden representations of the large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are constrained to be consistent with those of the large model, to transfer task-specific features which guarantee task performance. With G2SD, the vanilla ViT-Small model respectively achieves 98.7%, 98.1% and 99.3% the performance of its teacher (ViT-Base) for image classification, object detection, and semantic segmentation, setting a solid baseline for two-stage vision distillation. Code will be available at https://github.com/pengzhiliang/G2SD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0f7e73e-2553-47b8-bcc0-399aed4ac859Cited by top-tier papers8
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta et al.ICML 2024 · 15 citations
- SparseMAE: Sparse Training Meets Masked AutoencodersAojun Zhou, Yang Li, Zipeng Qin, Jianbo Liu et al.ICCV 2023 · 8 citations
- Asymmetric Masked Distillation for Pre-Training Small Foundation ModelsZhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu et al.CVPR 2024 · 6 citations
- Boosting Vanilla Lightweight Vision Transformers via Re-parameterizationZhentao Tan, Xiaodan Li, Yue Wu, Qi Chu et al.ICLR 2024 · 5 citations
- Mask Again: Masked Knowledge Distillation for Masked Video ModelingXiaojie Li, Shaowei He, Jianlong Wu, Yue Yu et al.ACM MM 2023 · 3 citations
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Generalizable Knowledge Distillation from Vision Foundation Models for Semantic SegmentationChonghua Lv, Dong Zhao, Shuang Wang, Dou Quan et al.CVPR 2026 · 1 citation
- Exploring Target Representations for Masked AutoencodersXingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin et al.ICLR 2024 · 59 citations
- TinyMIM: An Empirical Study of Distilling MIM Pre-trained ModelsSucheng Ren, Fangyun Wei, Zheng Zhang, Han HuCVPR 2023
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2023
- A Closer Look at Self-Supervised Lightweight Vision TransformersShaoru Wang, Jin Gao, Zeming Li, Xiaoqin Zhang et al.ICML 2023 · 61 citations
