The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture Model
Kaito Takanami, Takashi Takahashi, Ayaka Sakata
摘要
Self-distillation (SD), a technique where a model improves itself using its own predictions, has attracted attention as a simple yet powerful approach in machine learning. Despite its widespread use, the mechanisms underlying its effectiveness remain unclear. In this study, we investigate the efficacy of hyperparameter-tuned multi-stage SD with a linear classifier for binary classification on noisy Gaussian mixture data. For the analysis, we employ the replica method from statistical physics. Our findings reveal that the primary driver of SD's performance improvement is denoising through hard pseudo-labels, namely discrete labels generated from the model's own predictions, with the most notable gains observed in moderately sized datasets. We also identify two practical heuristics to enhance SD: early stopping that limits the number of stages, which is broadly effective, and bias parameter fixing, which helps under label imbalance. To empirically validate our theoretical findings derived from our toy model, we conduct additional experiments on CIFAR-10 classification using pretrained ResNet backbone. These results provide both theoretical and practical insights, advancing our understanding and application of SD in noisy settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active LearningHugo Cui, Yue LuICML 2026
- Self-Boost via Optimal Retraining: An Analysis via Approximate Message PassingAdel Javanmard, Rudrajit Das, Alessandro Epasto, Vahab MirrokniNeurIPS 2025
它引用的顶会 Paper15
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- The Role of Regularization in Classification of High-dimensional Noisy Gaussian MixtureFrancesca Mignacco, Florent Krzakala, Yue M. Lu, Pierfrancesco Urbani 等ICML 2020 · 被引用 98 次
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
相关 Paper
- Understanding Self-Distillation in the Presence of Label NoiseRudrajit Das, Sujay SanghaviICML 2023 · 被引用 25 次
- Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial LabelsHyeonsu Jeong, Hye Won ChungICLR 2025
- Self-cognitive Denoising in the Presence of Multiple Noisy Label SourcesYi-Xuan Sun, Ya-Lin Zhang, Bin Han, Longfei Li 等ICML 2024 · 被引用 2 次
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 被引用 1 次
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
