Debiased Distillation for Consistency Regularization
Lu Wang, Liuchi Xu, Xiong Yang, Zhenhua Huang, Jun Cheng
摘要
Knowledge distillation transfers "dark knowledge" from a large teacher model to a smaller student model, yielding a highly efficient network. To improve the network's generalization ability, existing works use a larger temperature coefficient for knowledge distillation. Nevertheless, these methods may reduce the confidence of the target category and lead to ambiguous recognition of similar samples. To mitigate this issue, some studies introduce intra-batch distillation to reduce prediction discrepancy. However, these methods overlook the inconsistency between background information and the target category, which may increase prediction bias due to noise disturbance. Additionally, label imbalance from random sampling and batch size can undermine network generalization reliability. To tackle these challenges, we propose a simple yet effective Intra-class Knowledge Distillation (IKD) method that facilitates knowledge sharing within the same class to ensure consistent predictions. First, we initialize the matrix and the vector to store logits and class counts provided by the teacher, respectively. Then, in the first epoch, we calculate the sum of logits and sample counts per class and perform KD to prevent knowledge omission. Finally, in subsequent training, we update the matrix to obtain the average logits and compute the KL divergence between the student's output and the updated matrix according to the label index. This process ensures intra-class consistency and improves the student's performance. Furthermore, this method theoretically reduces prediction bias by ensuring intra-class consistency. Extensive experiments on the CIFAR-100, ImageNet-1K, and Tiny-ImageNet datasets validate the superiority of IKD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Local Dense Logit Relations for Enhanced Knowledge DistillationLiuchi Xu, Kang Liu, Jinshuai Liu, Lu Wang 等ICCV 2025 · 被引用 10 次
- DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation TrainerHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenNeurIPS 2025 · 被引用 3 次
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 被引用 1 次
- Heterogeneous Complementary DistillationLiuchi Xu, Hao Zheng, Lu Wang, Lisheng Xu 等AAAI 2026
- Learning from the Undesirable: Robust Adaptation of Language Models Without ForgettingYunhun Nam, Jaehyung Kim, Jongheon JeongAAAI 2026
它引用的顶会 Paper34
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
相关 Paper
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
- Exploring Inter-Channel Correlation for Diversity-preserved Knowledge DistillationLi Liu, Qingle Huang, Sihao Lin, Hongwei Xie 等ICCV 2021 · 被引用 132 次
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 被引用 8 次
- Enhancing Class-Imbalanced Learning with Pre-Trained Guidance through Class-Conditional Knowledge DistillationLan Li, Xin-Chun Li, Han-Jia Ye, De-Chuan ZhanICML 2024 · 被引用 5 次
- Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental LearningZijian Gao, Shanhao Han, Xingxing Zhang, Kele Xu 等AAAI 2025 · 被引用 10 次
