Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across Domains
Haojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang, Yaliang Li, Jun Huang
摘要
Pre-trained language models have been applied to various NLP tasks with considerable performance gains. However, the large model sizes, together with the long inference time, limit the deployment of such models in realtime applications. One line of model compression approaches considers knowledge distillation to distill large teacher models into small student models. Most of these studies focus on single-domain only, which ignores the transferable knowledge from other domains. We notice that training a teacher with transferable knowledge digested across domains can achieve better generalization capability to help knowledge distillation. Hence we propose a Meta-Knowledge Distillation (Meta-KD) framework to build a meta-teacher model that captures transferable knowledge across domains and passes such knowledge to students. Specifically, we explicitly force the meta-teacher to capture transferable knowledge at both instance-level and feature-level from multiple domains, and then propose a meta-distillation algorithm to learn singledomain student models with guidance from the meta-teacher. Experiments on public multidomain NLP tasks show the effectiveness and superiority of the proposed Meta-KD framework. Further, we also demonstrate the capability of Meta-KD in the settings where the training data is scarce.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- BERT Learns to Teach: Knowledge Distillation with Meta LearningWangchunshu Zhou, Canwen Xu, Julian J. McAuleyACL 2022 · 被引用 114 次
- TransPrompt: Towards an Automatic Transferable Prompting Framework for Few-shot Text ClassificationChengyu Wang, Jianing Wang, Minghui Qiu, Jun Huang 等EMNLP 2021 · 被引用 39 次
- Search for Efficient Large Language ModelsXuan Shen, Pu Zhao, Yifan Gong, Zhenglun Kong 等NeurIPS 2024 · 被引用 23 次
- A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge DistillationAyan Sengupta, Shantanu Dixit, Md. Shad Akhtar, Tanmoy ChakrabortyICLR 2024 · 被引用 16 次
它引用的顶会 Paper10
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
- Zero-shot Text Classification via Reinforced Self-trainingZhiquan Ye, Yuxia Geng, Jiaoyan Chen, Jingmin Chen 等ACL 2020 · 被引用 75 次
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Meta-Learning Deep Energy-Based Memory ModelsSergey Bartunov, Jack W. Rae, Simon Osindero, Timothy P. LillicrapICLR 2020 · 被引用 35 次
相关 Paper
- HRKD: Hierarchical Relational Knowledge Distillation for Cross-domain Language Model CompressionChenhe Dong, Yaliang Li, Ying Shen, Minghui QiuEMNLP 2021 · 被引用 6 次
- Reinforced Multi-Teacher Selection for Knowledge DistillationFei Yuan, Linjun Shou, Jian Pei, Wutao Lin 等AAAI 2021 · 被引用 155 次
- Learning to Augment for Data-scarce Domain BERT Knowledge DistillationLingyun Feng, Minghui Qiu, Yaliang Li, Hai-Tao Zheng 等AAAI 2021 · 被引用 12 次
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 被引用 8 次
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
