Understanding Self-Distillation in the Presence of Label Noise
Rudrajit Das, Sujay Sanghavi
Abstract
Self-distillation (SD) is the process of first training a teacher model and then using its predictions to train a student model with the same architecture. Specifically, the student's objective function is , where is some loss function and is some parameter . Empirically, SD has been observed to provide performance gains in several settings. In this paper, we theoretically characterize the effect of SD in two supervised learning problems with noisy labels. We first analyze SD for regularized linear regression and show that in the high label noise regime, the optimal value of that minimizes the expected error in estimating the ground truth parameter is surprisingly greater than 1. Empirically, we show that works better than even with the cross-entropy loss for several classification datasets when 50% or 30% of the labels are corrupted. Further, we quantify when optimal SD is better than optimal regularization. Next, we analyze SD in the case of logistic regression for binary classification with random label corruption and quantify the range of label corruption in which the student outperforms the teacher in terms of accuracy. To our knowledge, this is the first result of its kind for the cross-entropy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58653c4d-d1ef-4300-bc85-df8f54ab449fCited by top-tier papers16
- Understanding the Gains from Repeated Self-DistillationDivyansh Pareek, Simon S. Du, Sewoong OhNeurIPS 2024 · 16 citations
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 8 citations
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton et al.ICLR 2025 · 6 citations
- GRAM-R²: Self-Training Generative Foundation Reward Models for Reward ReasoningChenglong Wang, Yongyu Mu, Hang Zhou, Yifu Huo et al.AAAI 2026 · 5 citations
- The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture ModelKaito Takanami, Takashi Takahashi, Ayaka SakataNeurIPS 2025 · 4 citations
Builds on12
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
Related papers
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 1 citation
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial LabelsHyeonsu Jeong, Hye Won ChungICLR 2025
- Even your Teacher Needs Guidance: Ground-Truth Targets Dampen Regularization Imposed by Self-DistillationKenneth Borup, Lars Nørvang AndersenNeurIPS 2021 · 18 citations
- Self-cognitive Denoising in the Presence of Multiple Noisy Label SourcesYi-Xuan Sun, Ya-Lin Zhang, Bin Han, Longfei Li et al.ICML 2024 · 2 citations
