Understanding Self-Distillation in the Presence of Label Noise
Rudrajit Das, Sujay Sanghavi
摘要
Self-distillation (SD) is the process of first training a teacher model and then using its predictions to train a student model with the same architecture. Specifically, the student's objective function is , where is some loss function and is some parameter . Empirically, SD has been observed to provide performance gains in several settings. In this paper, we theoretically characterize the effect of SD in two supervised learning problems with noisy labels. We first analyze SD for regularized linear regression and show that in the high label noise regime, the optimal value of that minimizes the expected error in estimating the ground truth parameter is surprisingly greater than 1. Empirically, we show that works better than even with the cross-entropy loss for several classification datasets when 50% or 30% of the labels are corrupted. Further, we quantify when optimal SD is better than optimal regularization. Next, we analyze SD in the case of logistic regression for binary classification with random label corruption and quantify the range of label corruption in which the student outperforms the teacher in terms of accuracy. To our knowledge, this is the first result of its kind for the cross-entropy loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Understanding the Gains from Repeated Self-DistillationDivyansh Pareek, Simon S. Du, Sewoong OhNeurIPS 2024 · 被引用 16 次
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 被引用 8 次
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton 等ICLR 2025 · 被引用 6 次
- GRAM-R²: Self-Training Generative Foundation Reward Models for Reward ReasoningChenglong Wang, Yongyu Mu, Hang Zhou, Yifu Huo 等AAAI 2026 · 被引用 5 次
- The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture ModelKaito Takanami, Takashi Takahashi, Ayaka SakataNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper12
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
相关 Paper
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 被引用 1 次
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 被引用 411 次
- Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial LabelsHyeonsu Jeong, Hye Won ChungICLR 2025
- Even your Teacher Needs Guidance: Ground-Truth Targets Dampen Regularization Imposed by Self-DistillationKenneth Borup, Lars Nørvang AndersenNeurIPS 2021 · 被引用 18 次
- Self-cognitive Denoising in the Presence of Multiple Noisy Label SourcesYi-Xuan Sun, Ya-Lin Zhang, Bin Han, Longfei Li 等ICML 2024 · 被引用 2 次
