Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm
Shaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang, Sung-En Chang, Bingbing Li, Shiyang Chen, Mimi Xie, Sanguthevar Rajasekaran, Hang Liu, Caiwen Ding
Abstract
Conventional wisdom in pruning Transformerbased language models is that pruning reduces the model expressiveness and thus is more likely to underfit rather than overfit. However, under the trending pretrain-and-finetune paradigm, we postulate a counter-traditional hypothesis, that is: pruning increases the risk of overfitting when performed at the fine-tuning phase. In this paper, we aim to address the overfitting problem and improve pruning performance via progressive knowledge distillation with error-bound properties. We show for the first time that reducing the risk of overfitting can help the effectiveness of pruning under the pretrain-and-finetune paradigm. Ablation studies and experiments on the GLUE benchmark show that our method outperforms the leading competitors across different tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7c14c3b-496d-494b-972b-ff7b1efa6fc6Cited by top-tier papers9
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 96 citations
- Effective Model Sparsification by Scheduled Grow-and-Prune MethodsXiaolong Ma, Minghai Qin, Fei Sun, Zejiang Hou et al.ICLR 2022 · 45 citations
- AutoReP: Automatic ReLU Replacement for Fast Private Network InferenceHongwu Peng, Shaoyi Huang, Tong Zhou, Yukui Luo et al.ICCV 2023 · 44 citations
- PROD: Progressive Distillation for Dense RetrievalZhenghao Lin, Yeyun Gong, Xiao Liu, Hang Zhang et al.WWW 2023 · 33 citations
- HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural NetworksMiao Yin, Yang Sui, Wanzhao Yang, Xiao Zang et al.CVPR 2022 · 17 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
Related papers
- Gradient-based Intra-attention Pruning on Pre-trained Language ModelsZiqing Yang, Yiming Cui, Xin Yao, Shijin WangACL 2023 · 2 citations
- From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model CompressionRunxin Xu, Fuli Luo, Chengyu Wang, Baobao Chang et al.AAAI 2022 · 32 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and InferenceBowen Zhao, Hannaneh Hajishirzi, Qingqing CaoICML 2024 · 31 citations
- Sparse Teachers Can Be Dense with KnowledgeYi Yang, Chen Zhang, Dawei SongEMNLP 2022 · 2 citations
