Instruction-Following Pruning for Large Language Models
Bairu Hou, Qibin Chen, Jianyu Wang, Guoli Yin, Chong Wang, Nan Du, Ruoming Pang, Shiyu Chang, Tao Lei
Abstract
With the rapid scaling of large language models (LLMs), structured pruning has become a widely used technique to learn efficient, smaller models from larger ones, delivering superior performance compared to training similarly sized models from scratch. In this paper, we move beyond the traditional static pruning approach of determining a fixed pruning mask for a model, and propose a dynamic approach to structured pruning. In our method, the pruning mask is inputdependent and adapts dynamically based on the information described in a user instruction. Our approach, termed "instruction-following pruning", introduces a sparse mask predictor that takes the user instruction as input and dynamically selects the most relevant model parameters for the given task. To identify and activate effective parameters, we jointly optimize the sparse mask predictor and the LLM, leveraging both instruction-following data and the pre-training corpus. Experimental results demonstrate the effectiveness of our approach on a wide range of evaluation benchmarks. For example, our 3B activated model improves over the 3B dense model by 5-8 points of absolute margin on domains such as math and coding, and rivals the performance of a 9B model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b618eb88-1b50-49ef-9ed3-9dccdd72b211Cited by top-tier papers6
- KVLink: Accelerating Large Language Models via Efficient KV Cache ReuseJingbo Yang, Bairu Hou, Wei Wei, Yujia Bao et al.NeurIPS 2025 · 83 citations
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof et al.ICLR 2026 · 11 citations
- SubspacePath Pruner: Inference-time Pruning via Probe-based Representation–Parameter CouplingZhiren Gong, Yikun Hou, Fan Wu, CHE WANG et al.ICML 2026
- Persona-Pruner: Sculpting Lightweight Models for Role-PlayingJinsu Kim, Jihoon Tack, Noah Lee, Jongheon JeongICML 2026
- Masked Gated Linear UnitYukito Tajima, Nakamasa Inoue, Yusuke Sekikawa, Ikuro Sato et al.NeurIPS 2025
Builds on15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
Related papers
- Let LLM Tell What to Prune and How Much to PruneMingzhe Yang, Sihao Lin, Changlin Li, Xiaojun ChangICML 2025
- Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and EfficiencyYanyang Li, Fuli Luo, Runxin Xu, Songfang Huang et al.ACL 2022 · 3 citations
- Pruning Pre-trained Language Models Without Fine-TuningTing Jiang, Deqing Wang, Fuzhen Zhuang, Ruobing Xie et al.ACL 2023 · 1 citation
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang et al.ICML 2025
- APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and InferenceBowen Zhao, Hannaneh Hajishirzi, Qingqing CaoICML 2024 · 31 citations
