MATE-KD: Masked Adversarial TExt, a Companion to Knowledge Distillation
Ahmad Rashid, Vasileios Lioutas, Mehdi Rezagholizadeh
Abstract
The advent of large pre-trained language models has given rise to rapid progress in the field of Natural Language Processing (NLP). While the performance of these models on standard benchmarks has scaled with size, compression techniques such as knowledge distillation have been key in making them practical. We present, MATE-KD, a novel textbased adversarial training algorithm which improves the performance of knowledge distillation. MATE-KD first trains a masked language model based generator to perturb text by maximizing the divergence between teacher and student logits. Then using knowledge distillation a student is trained on both the original and the perturbed training samples. We evaluate our algorithm, using BERT-based models, on the GLUE benchmark and demonstrate that MATE-KD outperforms competitive adversarial learning and data augmentation baselines. On the GLUE test set our 6 layer RoBERTa based model outperforms BERT LARGE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Towards Zero-Shot Knowledge Distillation for Natural Language ProcessingAhmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi RezagholizadehEMNLP 2021 · 25 citations
- Universal-KD: Attention-based Output-Grounded Intermediate Layer Knowledge DistillationYimeng Wu, Mehdi Rezagholizadeh, Abbas Ghaddar, Md. Akmal Haidar et al.EMNLP 2021 · 18 citations
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok et al.EMNLP 2022 · 4 citations
- Teacher Guided Training: An Efficient Framework for Knowledge TransferManzil Zaheer, Ankit Singh Rawat, Seungyeon Kim, Chong You et al.ICLR 2023 · 1 citation
Builds on3
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
Related papers
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 16 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen et al.EMNLP 2020 · 18 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
