Movement Pruning: Adaptive Sparsity by Fine-Tuning
Victor Sanh, Thomas Wolf, Alexander M. Rush
摘要
Magnitude pruning is a widely used strategy for reducing model size in pure supervised learning; however, it is less effective in the transfer learning regime that has become standard for state-of-the-art natural language processing applications. We propose the use of movement pruning, a simple, deterministic first-order weight pruning method that is more adaptive to pretrained model fine-tuning. We give mathematical foundations to the method and compare it to existing zeroth- and first-order pruning methods. Experiments show that when pruning large pretrained language models, movement pruning shows significant improvements in high-sparsity regimes. When combined with distillation, the approach achieves minimal accuracy loss with down to only 3% of the model parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper164
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Training with Quantization Noise for Extreme Model CompressionPierre Stock, Angela Fan, Benjamin Graham, Edouard Grave 等ICLR 2021 · 被引用 262 次
- Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLPHaonan Yu, Sergey Edunov, Yuandong Tian, Ari S. MorcosICLR 2020 · 被引用 156 次
相关 Paper
- Pruning Pre-trained Language Models Without Fine-TuningTing Jiang, Deqing Wang, Fuzhen Zhuang, Ruobing Xie 等ACL 2023 · 被引用 1 次
- Block Pruning For Faster TransformersFrançois Lagunas, Ella Charlaix, Victor Sanh, Alexander M. RushEMNLP 2021 · 被引用 2 次
- Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune ParadigmShaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang 等ACL 2022
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 被引用 236 次
- Gradient-based Intra-attention Pruning on Pre-trained Language ModelsZiqing Yang, Yiming Cui, Xin Yao, Shijin WangACL 2023 · 被引用 2 次
