Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
Qi Le, Enmao Diao, Ziyan Wang, Xinran Wang, Jie Ding, Li Yang, Ali Anwar
Abstract
We introduce Probe Pruning (PP), a novel framework for online, dynamic, structured pruning of Large Language Models (LLMs) applied in a batch-wise manner. PP leverages the insight that not all samples and tokens contribute equally to the model's output, and probing a small portion of each batch effectively identifies crucial weights, enabling tailored dynamic pruning for different batches. It comprises three main stages: probing, history-informed pruning, and full inference. In the probing stage, PP selects a small yet crucial set of hidden states, based on residual importance, to run a few model layers ahead. During the history-informed pruning stage, PP strategically integrates the probing states with historical states. Subsequently, it structurally prunes weights based on the integrated states and the PP importance score, a metric developed specifically to assess the importance of each weight channel in maintaining performance. In the final stage, full inference is conducted on the remaining weights. A major advantage of PP is its compatibility with existing models, as it operates without requiring additional neural network modules or fine-tuning. Comprehensive evaluations of PP on LLaMA-2/3 and OPT models reveal that even minimal probing-using just 1.5% of FLOPs-can substantially enhance the efficiency of structured pruning of LLMs. For instance, when evaluated on LLaMA-2-7B with WikiText2, PP achieves a 2.56× lower ratio of performance degradation per unit of runtime reduction compared to the state-of-the-art method at a 40% pruning ratio. Our code is available at https://github.com/Qi-Le1/Probe_Pruning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d92379a-0efb-48a7-8933-5b0da558d277Cited by top-tier papers9
- Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual SparsitySusav Shrestha, Bradley W. Settlemyer, Nikoli Dryden, A. L. Narasimha ReddyNeurIPS 2025 · 8 citations
- Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution OptimizationGuanchen Li, Yixing Xu, Zeping Li, Ji Liu et al.NeurIPS 2025 · 7 citations
- Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language ModelsMingge Lu, Jingwei Sun, Junqing Lin, Zechun Zhou et al.NeurIPS 2025 · 1 citation
- EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action ModelsYuting Huang, Leilei Ding, Zhipeng Tang, Zenghuan Zhu et al.ICML 2026 · 1 citation
- SubspacePath Pruner: Inference-time Pruning via Probe-based Representation–Parameter CouplingZhiren Gong, Yikun Hou, Fan Wu, CHE WANG et al.ICML 2026
Builds on22
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMsYang Ji, Ying SunACL 2026
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang et al.ICML 2025
- Bypass Back-propagation: Optimization-based Structural Pruning for Large Language Models via Policy GradientYuan Gao, Zujing Liu, Weizhong Zhang, Bo Du et al.ACL 2025
- Let LLM Tell What to Prune and How Much to PruneMingzhe Yang, Sihao Lin, Changlin Li, Xiaojun ChangICML 2025
- Dual-Assessment Driven Pruning: Iterative Optimizing Layer-wise Sparsity for Large Language ModelQinghui Sun, Weilun Wang, Yanni Zhu, Shenghuan He et al.KDD 2024 · 3 citations
