Self-Distillation Amplifies Regularization in Hilbert Space
Hossein Mobahi, Mehrdad Farajtabar, Peter L. Bartlett
摘要
Knowledge distillation introduced in the deep learning context is a method to transfer knowledge from one architecture to another. In particular, when the architectures are identical, this is called self-distillation. The idea is to feed in predictions of the trained model as new target values for retraining (and iterate this loop possibly a few times). It has been empirically observed that the self-distilled model often achieves higher accuracy on held out data. Why this happens, however, has been a mystery: the self-distillation dynamics does not receive any new information about the task and solely evolves by looping over training. To the best of our knowledge, there is no rigorous understanding of this phenomenon. This work provides the first theoretical analysis of self-distillation. We focus on fitting a nonlinear function to training data, where the model space is Hilbert space and fitting is subject to 2 regularization in this function space. We show that self-distillation iterations modify regularization by progressively limiting the number of basis functions that can be used to represent the solution. This implies (as we also verify empirically) that while a few rounds of self-distillation may reduce over-fitting, further rounds may lead to under-fitting and thus worse performance. * This article is a more detailed version of a paper with the same title in Neural and Information Processing Systems (NeurIPS) 2020 conference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper96
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang 等NeurIPS 2021 · 被引用 610 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 被引用 261 次
- Improve Object Detection with Feature-based Knowledge Distillation: Towards Accurate and Efficient DetectorsLinfeng Zhang, Kaisheng MaICLR 2021 · 被引用 251 次
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang 等CVPR 2022 · 被引用 213 次
它引用的顶会 Paper2
相关 Paper
- Even your Teacher Needs Guidance: Ground-Truth Targets Dampen Regularization Imposed by Self-DistillationKenneth Borup, Lars Nørvang AndersenNeurIPS 2021 · 被引用 18 次
- Understanding the Gains from Repeated Self-DistillationDivyansh Pareek, Simon S. Du, Sewoong OhNeurIPS 2024 · 被引用 16 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Understanding Self-Distillation in the Presence of Label NoiseRudrajit Das, Sujay SanghaviICML 2023 · 被引用 25 次
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
