Sparse Teachers Can Be Dense with Knowledge
Yi Yang, Chen Zhang, Dawei Song
摘要
Recent advances in distilling pretrained language models have discovered that, besides the expressiveness of knowledge, the student-friendliness should be taken into consideration to realize a truly knowledgeable teacher. Based on a pilot study, we find that over-parameterized teachers can produce expressive yet student-unfriendly knowledge and are thus limited in overall knowledgeableness. To remove the parameters that result in student-unfriendliness, we propose a sparse teacher trick under the guidance of an overall knowledgeable score for each teacher parameter. The knowledgeable score is essentially an interpolation of the expressiveness and student-friendliness scores. The aim is to ensure that the expressive parameters are retained while the student-unfriendly ones are removed. Extensive experiments on the GLUE benchmark show that the proposed sparse teachers can be dense with knowledge and lead to students with compelling performance in comparison with a series of competitive baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards the Law of Capacity Gap in Distilling Language ModelsChen Zhang, Qiuchi Li, Dawei Song, Zheyu Ye 等ACL 2025 · 被引用 39 次
- Lifting the Curse of Capacity Gap in Distilling Language ModelsChen Zhang, Yang Yang, Jiahao Liu, Jingang Wang 等ACL 2023 · 被引用 9 次
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 被引用 7 次
它引用的顶会 Paper9
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang 等NeurIPS 2020 · 被引用 401 次
相关 Paper
- In Good GRACES: Principled Teacher Selection for Knowledge DistillationAbhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Sham M. Kakade 等ICLR 2026 · 被引用 5 次
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok 等EMNLP 2022 · 被引用 4 次
- Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune ParadigmShaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang 等ACL 2022
- AD-KD: Attribution-Driven Knowledge Distillation for Language Model CompressionSiyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang 等ACL 2023 · 被引用 10 次
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey 等NeurIPS 2022 · 被引用 21 次
