Towards Zero-Shot Knowledge Distillation for Natural Language Processing
Ahmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi Rezagholizadeh
Abstract
Knowledge distillation (KD) is a common knowledge transfer algorithm used for model compression across a variety of deep learning based natural language processing (NLP) solutions. In its regular manifestations, KD requires access to the teacher's training data for knowledge transfer to the student network. However, privacy concerns, data regulations and proprietary reasons may prevent access to such data. We present, to the best of our knowledge, the first work on Zero-shot Knowledge Distillation for NLP, where the student learns from the much larger teacher without any task specific data. Our solution combines out-ofdomain data and adversarial training to learn the teacher's output distribution. We investigate six tasks from the GLUE benchmark and demonstrate that we can achieve between 75% and 92% of the teacher's classification score (accuracy or F1) while compressing the model 30 times.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9426313-a907-4532-89bb-185c1055439fCited by top-tier papers3
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu et al.EMNLP 2022 · 96 citations
- Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationYimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar et al.EMNLP 2021 · 18 citations
Builds on6
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot et al.ICLR 2020 · 244 citations
- ALP-KD: Attention-Based Layer Projection for Knowledge DistillationPeyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, Qun LiuAAAI 2021 · 142 citations
- Time-aware Large Kernel ConvolutionsVasileios Lioutas, Yuhong GuoICML 2020 · 30 citations
Related papers
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen et al.EMNLP 2020 · 18 citations
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 35 citations
- A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingNitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir KantorACL 2023 · 5 citations
