Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels
Hyeonsu Jeong, Hye Won Chung
Abstract
We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply. Our theoretical analysis reveals that multi-round self-distillation effectively performs label averaging among instances with high feature correlations, governed by the eigenvectors of the Gram matrix derived from input features. This process leads to clustered predictions and improved generalization, mitigating the impact of label noise by reducing the model's reliance on potentially corrupted labels. We establish conditions under which multi-round self-distillation achieves 100% population accuracy despite label noise. Furthermore, we introduce a novel, efficient single-round self-distillation method using refined partial labels from the teacher's top two softmax outputs, referred to as the PLL student model. This approach replicates the benefits of multiround distillation in a single round, achieving comparable or superior performanceespecially in high-noise scenarios-while significantly reducing computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37ee6a7a-0600-4542-9cda-8b2d60a56a62Cited by top-tier papers2
- Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data AnnotationMingxuan Xia, Haobo Wang, Yixuan Li, Zewei Yu et al.ACL 2025 · 4 citations
- Quantifying Cross-Domain Knowledge Distillation in the Presence of Domain ShiftXiangchao Li, Xiao Han, Qing Yang, Xin TongICML 2026
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- Progressive Identification of True Labels for Partial-Label LearningJiaqi Lv, Miao Xu, Lei Feng, Gang Niu et al.ICML 2020 · 220 citations
Related papers
- Understanding Self-Distillation in the Presence of Label NoiseRudrajit Das, Sujay SanghaviICML 2023 · 25 citations
- The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture ModelKaito Takanami, Takashi Takahashi, Ayaka SakataNeurIPS 2025 · 4 citations
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 155 citations
- Understanding the Gains from Repeated Self-DistillationDivyansh Pareek, Simon S. Du, Sewoong OhNeurIPS 2024 · 16 citations
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng et al.ICCV 2023 · 16 citations
