Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels
Hyeonsu Jeong, Hye Won Chung
摘要
We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply. Our theoretical analysis reveals that multi-round self-distillation effectively performs label averaging among instances with high feature correlations, governed by the eigenvectors of the Gram matrix derived from input features. This process leads to clustered predictions and improved generalization, mitigating the impact of label noise by reducing the model's reliance on potentially corrupted labels. We establish conditions under which multi-round self-distillation achieves 100% population accuracy despite label noise. Furthermore, we introduce a novel, efficient single-round self-distillation method using refined partial labels from the teacher's top two softmax outputs, referred to as the PLL student model. This approach replicates the benefits of multiround distillation in a single round, achieving comparable or superior performanceespecially in high-noise scenarios-while significantly reducing computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data AnnotationMingxuan Xia, Haobo Wang, Yixuan Li, Zewei Yu 等ACL 2025 · 被引用 4 次
- Quantifying Cross-Domain Knowledge Distillation in the Presence of Domain ShiftXiangchao Li, Xiao Han, Qing Yang, Xin TongICML 2026
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- Progressive Identification of True Labels for Partial-Label LearningJiaqi Lv, Miao Xu, Lei Feng, Gang Niu 等ICML 2020 · 被引用 220 次
相关 Paper
- Understanding Self-Distillation in the Presence of Label NoiseRudrajit Das, Sujay SanghaviICML 2023 · 被引用 25 次
- The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture ModelKaito Takanami, Takashi Takahashi, Ayaka SakataNeurIPS 2025 · 被引用 4 次
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 被引用 155 次
- Understanding the Gains from Repeated Self-DistillationDivyansh Pareek, Simon S. Du, Sewoong OhNeurIPS 2024 · 被引用 16 次
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng 等ICCV 2023 · 被引用 16 次
