Revisit the Power of Vanilla Knowledge Distillation: from Small Scale to Large Scale
Zhiwei Hao, Jianyuan Guo, Kai Han, Han Hu, Chang Xu, Yunhe Wang
Abstract
The tremendous success of large models trained on extensive datasets demonstrates that scale is a key ingredient in achieving superior results. Therefore, the reflection on the rationality of designing knowledge distillation (KD) approaches for limited-capacity architectures solely based on small-scale datasets is now deemed imperative. In this paper, we identify the small data pitfall that presents in previous KD methods, which results in the underestimation of the power of vanilla KD framework on large-scale datasets such as ImageNet-1K. Specifically, we show that employing stronger data augmentation techniques and using larger datasets can directly decrease the gap between vanilla KD and other meticulously designed KD variants. This highlights the necessity of designing and evaluating KD approaches in the context of practical scenarios, casting off the limitations of small-scale datasets. Our investigation of the vanilla KD and its variants in more complex schemes, including stronger training strategies and different model capacities, demonstrates that vanilla KD is elegantly simple but astonishingly effective in large-scale scenarios. Without bells and whistles, we obtain state-of-the-art ResNet-50, ViT-S, and ConvNeXtV2-T models for ImageNet, which achieve 83.1%, 84.3%, and 85.0% top-1 accuracy, respectively. PyTorch code and checkpoints can be found at https://github.com/Hao840/vanillaKD . Thus far, most existing KD approaches in the literature are tailored for small-scale benchmarks (e.g., CIFAR [19] ) and small teacher-student pairs (e.g., Res34-Res18 [2] and WRN40-WRN16 [10]). However, downstream vision tasks [20, 21, 22] actually require the backbone models to be pre-trained on large-scale datasets (e.g., ImageNet [23]) to achieve state-of-the-art performances. Only exploring KD approaches on small-scale datasets may fall short in providing a comprehensive understanding * Equal contribution. † Corresponding author. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef93fc2a-98d8-4339-a53e-603e940aead5Cited by top-tier papers6
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 59 citations
- Relational Diffusion Distillation for Efficient Image GenerationWeilun Feng, Chuanguang Yang, Zhulin An, Libo Huang et al.ACM MM 2024 · 11 citations
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 7 citations
- Enhancing Vision Transformer: Amplifying Non-Linearity in Feedforward Network ModuleYixing Xu, Chao Li, Dong Li, Xiao Sheng et al.ICML 2024 · 5 citations
- Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-DistillersYongqi Ding, Lin Zuo, Mengmeng Jing, Kunshan Yang et al.NeurIPS 2025 · 4 citations
Builds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 20 citations
- Are Large Kernels Better Teachers than Transformers for ConvNets?Tianjin Huang, Lu Yin, Zhenyu Zhang, Li Shen et al.ICML 2023 · 18 citations
- Knowledge distillation: A good teacher is patient and consistentLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva et al.CVPR 2022 · 215 citations
- : Improving Knowledge Distillation Using Orthogonal ProjectionsRoy Miles, Ismail Elezi, Jiankang DengCVPR 2024 · 9 citations
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 8 citations
