Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
Chen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu, Pengcheng He, Tuo Zhao, Weizhu Chen
Abstract
The Lottery Ticket Hypothesis suggests that an over-parametrized network consists of "lottery tickets", and training a certain collection of them (i.e., a subnetwork) can match the performance of the full model. In this paper, we study such a collection of tickets, which is referred to as "winning tickets", in extremely over-parametrized models, e.g., pre-trained language models. We observe that at certain compression ratios, the generalization performance of the winning tickets can not only match but also exceed that of the full model. In particular, we observe a phase transition phenomenon: As the compression ratio increases, generalization performance of the winning tickets first improves then deteriorates after a certain threshold. We refer to the tickets on the threshold as "super tickets". We further show that the phase transition is task and model dependent -as the model size becomes larger and the training data set becomes smaller, the transition becomes more pronounced. Our experiments on the GLUE benchmark show that the super tickets improve single task fine-tuning by 0.9 points on BERT-base and 1.0 points on BERT-large, in terms of task-average score. We also demonstrate that adaptively sharing the super tickets across tasks benefits multi-task learning 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f567227-695b-4a83-8ef6-9ecbae0198ecCited by top-tier papers28
- LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse ApproximationYixiao Li, Yifan Yu, Qingru Zhang, Chen Liang et al.ICML 2023 · 125 citations
- PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight ImportanceQingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin et al.ICML 2022 · 107 citations
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 100 citations
- From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model CompressionRunxin Xu, Fuli Luo, Chengyu Wang, Baobao Chang et al.AAAI 2022 · 32 citations
- Adaptive Budget Allocation for Parameter-Efficient Fine-TuningQingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He et al.ICLR 2023 · 32 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
Related papers
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen et al.ICML 2021 · 34 citations
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 114 citations
- EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan et al.ACL 2021
- Playing Lottery Tickets with Vision and LanguageZhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen et al.AAAI 2022 · 64 citations
- Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLPHaonan Yu, Sergey Edunov, Yuandong Tian, Ari S. MorcosICLR 2020 · 156 citations
