Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?
Keshigeyan Chandrasegaran, Ngoc-Trung Tran, Yunqing Zhao, Ngai-Man Cheung
Abstract
This work investigates the compatibility between label smoothing (LS) and knowledge distillation (KD). Contemporary findings addressing this thesis statement take dichotomous standpoints: Muller et al. (2019) and Shen et al. (2021b). Critically, there is no effort to understand and resolve these contradictory findings, leaving the primal question -- to smooth or not to smooth a teacher network? -- unanswered. The main contributions of our work are the discovery, analysis and validation of systematic diffusion as the missing concept which is instrumental in understanding and resolving these contradictory findings. This systematic diffusion essentially curtails the benefits of distilling from an LS-trained teacher, thereby rendering KD at increased temperatures ineffective. Our discovery is comprehensively supported by large-scale experiments, analyses and case studies including image classification, neural machine translation and compact student distillation tasks spanning across multiple datasets and teacher-student architectures. Based on our analysis, we suggest practitioners to use an LS-trained teacher with a low-temperature transfer to achieve high performance students. Code and models are available at https://keshik6.github.io/revisiting-ls-kd-compatibility/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Curriculum Temperature for Knowledge DistillationZheng Li, Xiang Li, Lingfeng Yang, Borui Zhao et al.AAAI 2023 · 277 citations
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang et al.CVPR 2024 · 183 citations
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 103 citations
- What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical PerspectiveHuan Wang, Suhas Lohit, Michael J. Jones, Yun FuNeurIPS 2022 · 62 citations
- Few-shot Image Generation via Adaptation-Aware Kernel ModulationYunqing Zhao, Keshigeyan Chandrasegaran, Milad Abdollahzadeh, Ngai-Man CheungNeurIPS 2022 · 55 citations
Builds on16
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- Fourier Spectrum Discrepancies in Deep Network Generated ImagesTarik Dzanic, Karan Shah, Freddie D. WitherdenNeurIPS 2020 · 235 citations
- Few-Shot Image Recognition With Knowledge TransferZhimao Peng, Zechao Li, Junge Zhang, Yan Li et al.ICCV 2019 · 230 citations
Related papers
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical StudyZhiqiang Shen, Zechun Liu, Dejia Xu, Zitian Chen et al.ICLR 2021 · 83 citations
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang et al.CVPR 2020
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 155 citations
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li et al.NeurIPS 2022 · 46 citations
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You et al.NeurIPS 2023 · 125 citations
