Can Students Beyond the Teacher? Distilling Knowledge from Teacher's Bias
Jianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng, Houxiang Zhang, Shengyong Chen
摘要
Knowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fundamental issue affecting student performance is the bias transferred by the teacher. Current KD frameworks transmit both right and wrong knowledge, introducing bias that misleads the student model. To address this issue, we propose a novel strategy to rectify bias and greatly improve the student model's performance. Our strategy involves three steps: First, we differentiate knowledge and design a bias elimination method to filter out biases, retaining only the right knowledge for the student model to learn. Next, we propose a bias rectification method to rectify the teacher model's wrong predictions, fundamentally addressing bias interference. The student model learns from both the right knowledge and the rectified biases, greatly improving its prediction accuracy. Additionally, we introduce a dynamic learning approach with a loss function that updates weights dynamically, allowing the student model to quickly learn right knowledge-based easy tasks initially and tackle hard tasks corresponding to biases later, greatly enhancing the student model's learning efficiency. To the best of our knowledge, this is the first strategy enabling the student model to surpass the teacher model. Experiments demonstrate that our strategy, as a plug-and-play module, is versatile across various mainstream KD frameworks. We will release our code after the paper is accepted.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Knowledge Distillation for Large Language Models through Residual LearningThinh On, Hengzhi Pei, Leonard Lausen, George KarypisICLR 2026 · 被引用 5 次
- Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identificationAihua Zheng, Pengyu Li, Zi Wang, Jin TangAAAI 2026
- FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance For Pruned Large Language ModelsJunyoung Lee, Sehyeon Park, Shinhyoung Jang, Seonha Ryu 等ICML 2026
它引用的顶会 Paper8
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- Curriculum Temperature for Knowledge DistillationZheng Li, Xiang Li, Lingfeng Yang, Borui Zhao 等AAAI 2023 · 被引用 277 次
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim 等NeurIPS 2021 · 被引用 134 次
相关 Paper
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
- DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language ModelsChangyi He, Yifu Ding, Jinyang Guo, Ruihao Gong 等ICML 2025
- Student-Oriented Teacher Knowledge Refinement for Knowledge DistillationChaomin Shen, Yaomin Huang, Haokun Zhu, Jinsong Fan 等ACM MM 2024 · 被引用 2 次
- Dynamic Knowledge Distillation for Pre-trained Language ModelsLei Li, Yankai Lin, Shuhuai Ren, Peng Li 等EMNLP 2021 · 被引用 32 次
- Comprehensive Knowledge Distillation with Causal InterventionXiang Deng, Zhongfei ZhangNeurIPS 2021 · 被引用 44 次
