Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers
Minjia Zhang, Uma-Naresh Niranjan, Yuxiong He
摘要
Deep and large pre-trained language models (e.g., BERT, GPT-3) are state-of-the-art for various natural language processing tasks. However, the huge size of these models brings challenges to fine-tuning and online deployment due to latency and cost constraints. Existing knowledge distillation methods reduce the model size, but they may encounter difficulties transferring knowledge from the teacher model to the student model due to the limited data from the downstream tasks. In this work, we propose AD^2, a novel and effective data augmentation approach to improving the task-specific knowledge transfer when compressing large pre-trained transformer models. Different from prior methods, AD^2 performs distillation by using an enhanced training set that contains both original inputs and adversarially perturbed samples that mimic the output distribution from the teacher.
Experimental results show that this method allows better transfer of knowledge from the teacher to the student during distillation, producing student models that retain 99.6% accuracy of the teacher model while outperforming existing task-specific knowledge distillation baselines by 1.2 points on average over a variety of natural language understanding tasks. Moreover, compared with alternative data augmentation methods, such as text-editing-based approaches, AD^2 is up to 28 times faster while achieving comparable or higher accuracy. In addition, when AD^2 is combined with more advanced task-agnostic distillation, we can advance the state-of-the-art performance even more. On top of the encouraging performance, this paper also provides thorough ablation studies and analysis. The discovered interplay between KD and adversarial data augmentation for compressing pre-trained Transformers may further inspire more advanced KD algorithms for compressing even larger scale models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok 等EMNLP 2022 · 被引用 4 次
- Augmentation with Projection: Towards an Effective and Efficient Data Augmentation Paradigm for DistillationZiqi Wang, Yuexin Wu, Frederick Liu, Daogao Liu 等ICLR 2023 · 被引用 2 次
- XTC: Extreme Compression for Pre-trained Transformers Made Simple and EfficientXiaoxia Wu, Zhewei Yao, Minjia Zhang, Conglong Li 等NeurIPS 2022 · 被引用 1 次
- Enhancing In-Context Learning via Implicit Demonstration AugmentationXiaoling Zhou, Wei Ye, Yidong Wang, Chaoya Jiang 等ACL 2024
- From Teacher Pathways to Invariant Manifolds: Consensus Subspace Distillation for TSFMsZexing Zhang, Tianyang Lei, Jichao Li, Yang KeweiICML 2026
它引用的顶会 Paper9
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun 等ICLR 2020 · 被引用 502 次
相关 Paper
- MATE-KD: Masked Adversarial TExt, a Companion to Knowledge DistillationAhmad Rashid, Vasileios Lioutas, Mehdi RezagholizadehACL 2021
- Learning to Augment for Data-scarce Domain BERT Knowledge DistillationLingyun Feng, Minghui Qiu, Yaliang Li, Hai-Tao Zheng 等AAAI 2021 · 被引用 12 次
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 被引用 8 次
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen 等EMNLP 2020 · 被引用 18 次
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
