Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
Seungcheol Park, Hojun Choi, U Kang
Abstract
Given a pretrained encoder-based language model, how can we accurately compress it without retraining? Retraining-free structured pruning algorithms are crucial in pretrained language model compression due to their significantly reduced pruning cost and capability to prune large language models. However, existing retraining-free algorithms encounter severe accuracy degradation, as they fail to handle pruning errors, especially at high compression rates. In this paper, we propose K-prune (Knowledge-preserving pruning), an accurate retraining-free structured pruning algorithm for pretrained encoder-based language models. K-prune focuses on preserving the useful knowledge of the pretrained model to minimize pruning errors through a carefully designed iterative pruning process composed of knowledge measurement, knowledge-preserving mask search, and knowledge-preserving weight-tuning. As a result, K-prune shows significant accuracy improvements up to 58.02%p higher F1 score compared to existing retraining-free pruning algorithms under a high compression rate of 80% on the SQuAD benchmark without any retraining process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12c67034-4e3d-44e6-86f5-3fbb507d1a04Cited by top-tier papers8
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim et al.ICLR 2026 · 5 citations
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient TransformersFiras Gabetni, Giuseppe Curci, Andrea Pilzer, Subhankar Roy et al.ICLR 2026 · 5 citations
- Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language ModelsSeungcheol Park, Jeongin Bae, Beomseok Kwon, Minjun Kim et al.ACL 2025 · 1 citation
- LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision TransformersMinjun Kim, Jaeri Lee, Jongjin Kim, Jeongin Yun et al.AAAI 2026 · 1 citation
- SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuningMinjun Kim, Jongjin Kim, U KangICLR 2025
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
Related papers
- Structured Optimal Brain Pruning for Large Language ModelsJiateng Wei, Quan Lu, Ning Jiang, Siqi Li et al.EMNLP 2024 · 2 citations
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun et al.NeurIPS 2022 · 247 citations
- Two-Stage Regularization-Based Structured Pruning for LLMsMingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang et al.ACL 2026 · 3 citations
- Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang et al.AAAI 2024 · 130 citations
- Gradient-Free Structured Pruning with Unlabeled DataAzade Nova, Hanjun Dai, Dale SchuurmansICML 2023 · 38 citations
