Knowledge Distillation from Internal Representations
Gustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Z. Yao, Xing Fan, Chenlei Guo
摘要
Knowledge distillation is typically conducted by training a small model (the student) to mimic a large and cumbersome model (the teacher). The idea is to compress the knowledge from the teacher by using its output probabilities as soft-labels to optimize the student. However, when the teacher is considerably large, there is no guarantee that the internal knowledge of the teacher will be transferred into the student; even if the student closely matches the soft-labels, its internal representations may be considerably different. This internal mismatch can undermine the generalization capabilities originally intended to be transferred from the teacher to the student. In this paper, we propose to distill the internal representations of a large model such as BERT into a simplified version of it. We formulate two ways to distill such representations and various algorithms to conduct the distillation. We experiment with datasets from the GLUE benchmark and consistently show that adding knowledge distillation from internal representations is a more powerful method than only using soft-label distillation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Zen-NAS: A Zero-Shot NAS for High-Performance Image RecognitionMing Lin, Pichao Wang, Zhenhong Sun, Hesen Chen 等ICCV 2021 · 被引用 164 次
- ALP-KD: Attention-Based Layer Projection for Knowledge DistillationPeyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, Qun LiuAAAI 2021 · 被引用 142 次
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 被引用 96 次
- BiT: Robustly Binarized Multi-distilled TransformerZechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao 等NeurIPS 2022 · 被引用 93 次
相关 Paper
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
- SKDBERT: Compressing BERT via Stochastic Knowledge DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2023 · 被引用 13 次
- Contrastive Distillation on Intermediate Representations for Language Model CompressionSiqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng 等EMNLP 2020 · 被引用 59 次
- Multi-Granularity Structural Knowledge Distillation for Language Model CompressionChang Liu, Chongyang Tao, Jiazhan Feng, Dongyan ZhaoACL 2022 · 被引用 64 次
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 被引用 8 次
