SLaM: Student-Label Mixing for Distillation with Unlabeled Examples
Vasilis Kontonis, Fotis Iliopoulos, Khoa Trinh, Cenk Baykal, Gaurav Menghani, Erik Vee
Abstract
Knowledge distillation with unlabeled examples is a powerful training paradigm for generating compact and lightweight student models in applications where the amount of labeled data is limited but one has access to a large pool of unlabeled data. In this setting, a large teacher model generates "soft" pseudo-labels for the unlabeled dataset which are then used for training the student model. Despite its success in a wide variety of applications, a shortcoming of this approach is that the teacher's pseudo-labels are often noisy, leading to impaired student performance. In this paper, we present a principled method for knowledge distillation with unlabeled examples that we call Student-Label Mixing (SLaM) and we show that it consistently improves over prior approaches by evaluating it on several standard benchmarks. Finally, we show that SLaM comes with theoretical guarantees; along the way we give an algorithm improving the best-known sample complexity for learning halfspaces with margin under random classification noise, and provide the first convergence analysis for so-called "forward loss-adjustment" methods. performance, and this is a well-known phenomenon that has been observed and studied in a plethora of papers in the literature, e.g., [6, 36, 44, 51, 53, 8, 27] . In this work, we propose Student-Label Mixing (SLaM), a principled method for knowledge distillation with unlabeled examples that accounts for the teacher's noise and consistently improves over prior approaches. At the heart of our method lies the observation that the noise introduced by the teacher is neither random nor adversarial, in the sense that it correlates well with metrics of "confidence" such as the margin score or the entropy of the teacher's predictions. We exploit this empirical fact to our benefit in order to introduce a model for the teacher's noise, which we use to appropriately modify the student's loss function. At a high level, for any given example during the student's training process, we evaluate the student's loss function on a convex combination of the student's current prediction and another (soft-)label that we estimate using our model for the teacher's noise (hence the name "student-label mixing"). Our contributions can be summarized as follows: 1. We propose SLaM: a principled method for improving knowledge distillation with unlabeled examples. The method is efficient, data-agnostic and simple to implement. 2. We provide extensive experimental evidence and comparisons which show that our method consistently outperforms previous approaches on standard benchmarks. Moreover, we show that SLaM can be combined with standard distillation techniques such as temperature scaling and confidence-based weighting schemes. 3. We give theoretical guarantees for SLaM under standard assumptions. As a byproduct of our analysis we obtain a simple "forward loss-adjustment" iteration that provably learns halfspaces with γ-margin under Random Classification Noise with O(1/(ϵ 2 γ 2 )) samples improving over prior works that had worse dependence on either the margin γ or the generalization error ϵ (see Theorem 5.1 and Remark 5.2). Related Work Knowledge Distillation. Most of the literature on knowledge distillation has been focused on the fully supervised/labeled setting, i.e., when distillation is performed on the labeled training data of the teacher model rather than on new, unlabeled data -see e.g. the original paper of [26] . Naturally, in this setting the pseudo-labels generated by the teacher are almost always accurate and so many follow-up works [2, 14, 15, 41, 52] have developed advanced distillation techniques that aim to enforce greater consistency between the teacher's and the student's predictions, or even between the intermediate representations learned by the two models. Applying such methods in our setting where the training dataset contains mainly unlabeled examples is still possible but, in this case, it is known [51, 27] that fully trusting the teacher model can be actually harmful to the student model, making these methods less effective. (In fact, when the teacher is highly noisy these methods even underperform vanilla distillation with unlabeled examples.) In Section 4.2 we present results that show the improved effectiveness of SLaM relative to the state-of-the-art supervised knowledge distillation methods like the Variational Information Distillation for Knowledge Transfer (VID) framework [2] . Moreover, in Appendix D.5 we show that our method can be combined with (i.e., provide an additional improvement) the most simple, yet surprisingly effective, methods of improving knowledge distillation, namely the temperature-scaling idea introduced by [26]. For distillation with unlabeled examples, many approaches [17, 33, 29] propose filtering-out or reweighting the teacher's pseudo-labels based on measures of teacher's uncertainty, such as dropout variance, entropy, margin-score, or the cut-statistic. These methods are i
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15eb087d-700d-4967-b0df-b5a5096bf4cbCited by top-tier papers4
- A Near-optimal Algorithm for Learning Margin Halfspaces with Massart NoiseIlias Diakonikolas, Nikos ZarifisNeurIPS 2024 · 8 citations
- Learning Noisy Halfspaces with a Margin: Massart is No Harder than RandomGautam Chandrasekaran, Vasilis Kontonis, Konstantinos Stavropoulos, Kevin TianNeurIPS 2024 · 8 citations
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy et al.NeurIPS 2025 · 5 citations
- Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence GenerationJiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer et al.ACL 2024
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
Related papers
- Weighted Distillation with Unlabeled ExamplesFotis Iliopoulos, Vasilis Kontonis, Cenk Baykal, Gaurav Menghani et al.NeurIPS 2022 · 20 citations
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 4 citations
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li et al.CVPR 2022 · 88 citations
- DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language ModelsChangyi He, Yifu Ding, Jinyang Guo, Ruihao Gong et al.ICML 2025
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian et al.ICML 2021 · 287 citations
