Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers
Minjia Zhang, Uma-Naresh Niranjan, Yuxiong He
Abstract
Deep and large pre-trained language models (e.g., BERT, GPT-3) are state-of-the-art for various natural language processing tasks. However, the huge size of these models brings challenges to fine-tuning and online deployment due to latency and cost constraints. Existing knowledge distillation methods reduce the model size, but they may encounter difficulties transferring knowledge from the teacher model to the student model due to the limited data from the downstream tasks. In this work, we propose AD^2, a novel and effective data augmentation approach to improving the task-specific knowledge transfer when compressing large pre-trained transformer models. Different from prior methods, AD^2 performs distillation by using an enhanced training set that contains both original inputs and adversarially perturbed samples that mimic the output distribution from the teacher.
Experimental results show that this method allows better transfer of knowledge from the teacher to the student during distillation, producing student models that retain 99.6% accuracy of the teacher model while outperforming existing task-specific knowledge distillation baselines by 1.2 points on average over a variety of natural language understanding tasks. Moreover, compared with alternative data augmentation methods, such as text-editing-based approaches, AD^2 is up to 28 times faster while achieving comparable or higher accuracy. In addition, when AD^2 is combined with more advanced task-agnostic distillation, we can advance the state-of-the-art performance even more. On top of the encouraging performance, this paper also provides thorough ablation studies and analysis. The discovered interplay between KD and adversarial data augmentation for compressing pre-trained Transformers may further inspire more advanced KD algorithms for compressing even larger scale models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok et al.EMNLP 2022 · 4 citations
- Augmentation with Projection: Towards an Effective and Efficient Data Augmentation Paradigm for DistillationZiqi Wang, Yuexin Wu, Frederick Liu, Daogao Liu et al.ICLR 2023 · 2 citations
- XTC: Extreme Compression for Pre-trained Transformers Made Simple and EfficientXiaoxia Wu, Zhewei Yao, Minjia Zhang, Conglong Li et al.NeurIPS 2022 · 1 citation
- Enhancing In-Context Learning via Implicit Demonstration AugmentationXiaoling Zhou, Wei Ye, Yidong Wang, Chaoya Jiang et al.ACL 2024
- From Teacher Pathways to Invariant Manifolds: Consensus Subspace Distillation for TSFMsZexing Zhang, Tianyang Lei, Jichao Li, Yang KeweiICML 2026
Builds on9
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
Related papers
- MATE-KD: Masked Adversarial TExt, a Companion to Knowledge DistillationAhmad Rashid, Vasileios Lioutas, Mehdi RezagholizadehACL 2021
- Learning to Augment for Data-scarce Domain BERT Knowledge DistillationLingyun Feng, Minghui Qiu, Yaliang Li, Hai-Tao Zheng et al.AAAI 2021 · 12 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen et al.EMNLP 2020 · 18 citations
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li et al.ACL 2025
