Revisit the Power of Vanilla Knowledge Distillation: from Small Scale to Large Scale
Zhiwei Hao, Jianyuan Guo, Kai Han, Han Hu, Chang Xu, Yunhe Wang
摘要
The tremendous success of large models trained on extensive datasets demonstrates that scale is a key ingredient in achieving superior results. Therefore, the reflection on the rationality of designing knowledge distillation (KD) approaches for limited-capacity architectures solely based on small-scale datasets is now deemed imperative. In this paper, we identify the small data pitfall that presents in previous KD methods, which results in the underestimation of the power of vanilla KD framework on large-scale datasets such as ImageNet-1K. Specifically, we show that employing stronger data augmentation techniques and using larger datasets can directly decrease the gap between vanilla KD and other meticulously designed KD variants. This highlights the necessity of designing and evaluating KD approaches in the context of practical scenarios, casting off the limitations of small-scale datasets. Our investigation of the vanilla KD and its variants in more complex schemes, including stronger training strategies and different model capacities, demonstrates that vanilla KD is elegantly simple but astonishingly effective in large-scale scenarios. Without bells and whistles, we obtain state-of-the-art ResNet-50, ViT-S, and ConvNeXtV2-T models for ImageNet, which achieve 83.1%, 84.3%, and 85.0% top-1 accuracy, respectively. PyTorch code and checkpoints can be found at https://github.com/Hao840/vanillaKD . Thus far, most existing KD approaches in the literature are tailored for small-scale benchmarks (e.g., CIFAR [19] ) and small teacher-student pairs (e.g., Res34-Res18 [2] and WRN40-WRN16 [10]). However, downstream vision tasks [20, 21, 22] actually require the backbone models to be pre-trained on large-scale datasets (e.g., ImageNet [23]) to achieve state-of-the-art performances. Only exploring KD approaches on small-scale datasets may fall short in providing a comprehensive understanding * Equal contribution. † Corresponding author. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 被引用 59 次
- Relational Diffusion Distillation for Efficient Image GenerationWeilun Feng, Chuanguang Yang, Zhulin An, Libo Huang 等ACM MM 2024 · 被引用 11 次
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 被引用 7 次
- Enhancing Vision Transformer: Amplifying Non-Linearity in Feedforward Network ModuleYixing Xu, Chao Li, Dong Li, Xiao Sheng 等ICML 2024 · 被引用 5 次
- Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-DistillersYongqi Ding, Lin Zuo, Mengmeng Jing, Kunshan Yang 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
相关 Paper
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 被引用 20 次
- Are Large Kernels Better Teachers than Transformers for ConvNets?Tianjin Huang, Lu Yin, Zhenyu Zhang, Li Shen 等ICML 2023 · 被引用 18 次
- Knowledge distillation: A good teacher is patient and consistentLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva 等CVPR 2022 · 被引用 215 次
- : Improving Knowledge Distillation Using Orthogonal ProjectionsRoy Miles, Ismail Elezi, Jiankang DengCVPR 2024 · 被引用 9 次
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 被引用 8 次
