Efficient Knowledge Distillation from Model Checkpoints
Chaofei Wang, Qisen Yang, Rui Huang, Shiji Song, Gao Huang
摘要
Knowledge distillation is an effective approach to learn compact models (students) with the supervision of large and strong models (teachers). As empirically there exists a strong correlation between the performance of teacher and student models, it is commonly believed that a high performing teacher is preferred. Consequently, practitioners tend to use a well trained network or an ensemble of them as the teacher. In this paper, we observe that an intermediate model, i.e., a checkpoint in the middle of the training procedure, often serves as a better teacher compared to the fully converged model, although the former has much lower accuracy. More surprisingly, a weak snapshot ensemble of several intermediate models from a same training trajectory can outperform a strong ensemble of independently trained and fully converged models, when they are used as teachers. We show that this phenomenon can be partially explained by the information bottleneck principle: the feature representations of intermediate models can have higher mutual information regarding the input, and thus contain more "dark knowledge" for effective distillation. We further propose an optimal intermediate teacher selection algorithm based on maximizing the total task-related mutual information. Experiments verify its effectiveness and applicability. Our code is available at https://github.com/LeapLabTHU/CheckpointKD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan 等NeurIPS 2023 · 被引用 138 次
- Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RLYang Yue, Rui Lu, Bingyi Kang, Shiji Song 等NeurIPS 2023 · 被引用 28 次
- Bayes Conditional Distribution Estimation for Knowledge Distillation Based on Conditional Mutual InformationLinfeng Ye, Shayan Mohajer Hamidi, Renhao Tan, En-Hui YangICLR 2024 · 被引用 27 次
- Trans-LoRA: towards data-free Transferable Parameter Efficient FinetuningRunqian Wang, Soumya Ghosh, David D. Cox, Diego Antognini 等NeurIPS 2024 · 被引用 15 次
- Fed-DFA: Federated Distillation for Heterogeneous Model Fusion Through the Adversarial LensZichen Wang, Feng Yan, Tianyi Wang, Cong Wang 等AAAI 2025 · 被引用 9 次
它引用的顶会 Paper10
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
相关 Paper
- Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge DistillationMartin Zong, Zengyu Qiu, Xinzhu Ma, Kunlin Yang 等ICLR 2023 · 被引用 19 次
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim 等NeurIPS 2021 · 被引用 134 次
- CIFD: Controlled Information Flow to Enhance Knowledge DistillationYashas Malur Saidutta, Rakshith Sharma Srinivasa, Jaejin Cho, Ching Hua Lee 等NeurIPS 2024 · 被引用 2 次
- Anomaly Detection via Reverse Distillation from One-Class EmbeddingHanqiu Deng, Xingyu LiCVPR 2022 · 被引用 701 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
