Gradient-based Intra-attention Pruning on Pre-trained Language Models
Ziqing Yang, Yiming Cui, Xin Yao, Shijin Wang
Abstract
Pre-trained language models achieve superior performance but are computationally expensive. Techniques such as pruning and knowledge distillation have been developed to reduce their sizes and latencies. In this work, we propose a structured pruning method GRAIN (Gradientbased Intra-attention pruning), which performs task-specific pruning with knowledge distillation and yields highly effective models. Different from common approaches that prune each attention head as a whole, GRAIN inspects and prunes intra-attention structures, which greatly expands the structure search space and enables more flexible models. We also propose a gradient separation strategy that reduces the interference of distillation on pruning for a better combination of the two approaches. Experiments on GLUE, SQuAD, and CoNLL 2003 show that GRAIN notably outperforms other methods, especially in the high sparsity regime, and achieves 6 ∼ 7× speedups while maintaining 93% ∼ 99% performance. Under extreme compression where only 3% transformer weights remain, the pruned model is still competitive compared to larger models. 1 1 Code is available at https://github.com/airaria/ GRAIN . 2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu et al.ICML 2024 · 64 citations
- LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language ModelsGuangyan Li, Yongqiang Tang, Wensheng ZhangICML 2024 · 11 citations
- Elastic ViTs from Pretrained Models without RetrainingWalter Simoncini, Michael Dorkenwald, Tijmen Blankevoort, Cees G. M. Snoek et al.NeurIPS 2025 · 2 citations
- Neural Parameter Search for Slimmer Fine-Tuned Models and Better TransferGuodong Du, Zitao Fang, Jing Li, Junlin Li et al.ACL 2025
Builds on13
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- DynaBERT: Dynamic BERT with Adaptive Width and DepthLu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang et al.NeurIPS 2020 · 401 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
Related papers
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 236 citations
- Gradient-Free Structured Pruning with Unlabeled DataAzade Nova, Hanjun Dai, Dale SchuurmansICML 2023 · 38 citations
- Block Pruning For Faster TransformersFrançois Lagunas, Ella Charlaix, Victor Sanh, Alexander M. RushEMNLP 2021 · 2 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune ParadigmShaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang et al.ACL 2022
