Towards Zero-Shot Knowledge Distillation for Natural Language Processing
Ahmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi Rezagholizadeh
摘要
Knowledge distillation (KD) is a common knowledge transfer algorithm used for model compression across a variety of deep learning based natural language processing (NLP) solutions. In its regular manifestations, KD requires access to the teacher's training data for knowledge transfer to the student network. However, privacy concerns, data regulations and proprietary reasons may prevent access to such data. We present, to the best of our knowledge, the first work on Zero-shot Knowledge Distillation for NLP, where the student learns from the much larger teacher without any task specific data. Our solution combines out-ofdomain data and adversarial training to learn the teacher's output distribution. We investigate six tasks from the GLUE benchmark and demonstrate that we can achieve between 75% and 92% of the teacher's classification score (accuracy or F1) while compressing the model 30 times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
- Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationYimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar 等EMNLP 2021 · 被引用 18 次
它引用的顶会 Paper6
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot 等ICLR 2020 · 被引用 244 次
- ALP-KD: Attention-Based Layer Projection for Knowledge DistillationPeyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, Qun LiuAAAI 2021 · 被引用 142 次
- Time-aware Large Kernel ConvolutionsVasileios Lioutas, Yuhong GuoICML 2020 · 被引用 30 次
相关 Paper
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen 等EMNLP 2020 · 被引用 18 次
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 等ICLR 2021 · 被引用 90 次
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang 等ACL 2021
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 被引用 35 次
- A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingNitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir KantorACL 2023 · 被引用 5 次
