Adversarial Self-Supervised Data-Free Distillation for Text Classification
Xinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen, Chenghao Jia, Weiming Lu
Abstract
Large pre-trained transformer-based language models have achieved impressive results on a wide range of NLP tasks. In the past few years, Knowledge Distillation(KD) has become a popular paradigm to compress a computationally expensive model to a resource-efficient lightweight model. However, most KD algorithms, especially in NLP, rely on the accessibility of the original training dataset, which may be unavailable due to privacy issues. To tackle this problem, we propose a novel twostage data-free distillation method, named Adversarial self-Supervised Data-Free Distillation (AS-DFD), which is designed for compressing large-scale transformer-based models (e.g., BERT). To avoid text generation in discrete space, we introduce a Plug & Play Embedding Guessing method to craft pseudo embeddings from the teacher's hidden knowledge. Meanwhile, with a self-supervised module to quantify the student's ability, we adapt the difficulty of pseudo embeddings in an adversarial training manner. To the best of our knowledge, our framework is the first data-free distillation framework designed for NLP tasks. We verify the effectiveness of our method on several text classification datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- Up to 100x Faster Data-Free Knowledge DistillationGongfan Fang, Kanya Mo, Xinchao Wang, Jie Song et al.AAAI 2022 · 103 citations
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo ReplayKuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman et al.AAAI 2022 · 59 citations
- Few-Shot Class-Incremental Learning for Named Entity RecognitionRui Wang, Tong Yu, Handong Zhao, Sungchul Kim et al.ACL 2022 · 26 citations
- PAT: Pruning-Aware Tuning for Large Language ModelsYijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang et al.AAAI 2025 · 1 citation
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
Related papers
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 35 citations
- Towards Zero-Shot Knowledge Distillation for Natural Language ProcessingAhmad Rashid, Vasileios Lioutas, Abbas Ghaddar, Mehdi RezagholizadehEMNLP 2021 · 25 citations
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 16 citations
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge DistillationHyunjune Shin, Dong-Wan ChoiAAAI 2024 · 8 citations
- Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model CompressionZhiwei Hao, Yong Luo, Han Hu, Jianping An et al.ACM MM 2021 · 11 citations
