Efficient Knowledge Distillation from Model Checkpoints
Chaofei Wang, Qisen Yang, Rui Huang, Shiji Song, Gao Huang
Abstract
Knowledge distillation is an effective approach to learn compact models (students) with the supervision of large and strong models (teachers). As empirically there exists a strong correlation between the performance of teacher and student models, it is commonly believed that a high performing teacher is preferred. Consequently, practitioners tend to use a well trained network or an ensemble of them as the teacher. In this paper, we observe that an intermediate model, i.e., a checkpoint in the middle of the training procedure, often serves as a better teacher compared to the fully converged model, although the former has much lower accuracy. More surprisingly, a weak snapshot ensemble of several intermediate models from a same training trajectory can outperform a strong ensemble of independently trained and fully converged models, when they are used as teachers. We show that this phenomenon can be partially explained by the information bottleneck principle: the feature representations of intermediate models can have higher mutual information regarding the input, and thus contain more "dark knowledge" for effective distillation. We further propose an optimal intermediate teacher selection algorithm based on maximizing the total task-related mutual information. Experiments verify its effectiveness and applicability. Our code is available at https://github.com/LeapLabTHU/CheckpointKD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7473528-7f7a-4e62-a30d-5f61d9ed191bCited by top-tier papers21
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan et al.NeurIPS 2023 · 138 citations
- Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RLYang Yue, Rui Lu, Bingyi Kang, Shiji Song et al.NeurIPS 2023 · 28 citations
- Bayes Conditional Distribution Estimation for Knowledge Distillation Based on Conditional Mutual InformationLinfeng Ye, Shayan Mohajer Hamidi, Renhao Tan, En-Hui YangICLR 2024 · 27 citations
- Trans-LoRA: towards data-free Transferable Parameter Efficient FinetuningRunqian Wang, Soumya Ghosh, David D. Cox, Diego Antognini et al.NeurIPS 2024 · 15 citations
- Fed-DFA: Federated Distillation for Heterogeneous Model Fusion Through the Adversarial LensZichen Wang, Feng Yan, Tianyi Wang, Cong Wang et al.AAAI 2025 · 9 citations
Builds on10
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
Related papers
- Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge DistillationMartin Zong, Zengyu Qiu, Xinzhu Ma, Kunlin Yang et al.ICLR 2023 · 19 citations
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim et al.NeurIPS 2021 · 134 citations
- CIFD: Controlled Information Flow to Enhance Knowledge DistillationYashas Malur Saidutta, Rakshith Sharma Srinivasa, Jaejin Cho, Ching Hua Lee et al.NeurIPS 2024 · 2 citations
- Anomaly Detection via Reverse Distillation from One-Class EmbeddingHanqiu Deng, Xingyu LiCVPR 2022 · 701 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
