OCP: Outlier-Centric Probing for Dynamic Structured Pruning of LLMs
Yang Ji, Ying Sun
Abstract
Structured pruning offers a hardware-friendly approach for efficient LLM inference. Early static methods determine fixed subnetworks through offline calibration, suffering from performance degradation and calibration sensitivity. Recent methods explore input-adaptive pruning by selecting a subset of tokens as probes to estimate hidden activations for online pruning decisions. However, existing probe selection strategies fail to identify outlier-triggering tokens, and uniform layerwise sparsity misaligns with heterogeneous outlier distributions, leading to critical channels being incorrectly pruned. Therefore, we propose OCP (Outlier-Centric Probing for structured pruning), a principled framework that prioritizes capturing outlier-triggering tokens rather than reconstructing full hidden distributions. Specifically, OCP includes three key components: (1) sensitivity-weighted probing for FFN layers that identifies outlier patterns via precomputed weight aggregations, (2) attention-accumulated probing that leverages preceding attention matrices to identify salient tokens, and (3) online adaptive sparsity allocation that dynamically adjusts layerwise pruning based on history-guided outlier statistics. Extensive experiments on LLaMA2, LLaMA3, and OPT demonstrate that OCP consistently outperforms state-of-the-art methods across benchmarks, achieving up to 25% perplexity reduction at 1.6× speedup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a91847a7-dca5-4f0f-9f26-941ed6af0790Builds on18
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High SparsityLu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh et al.ICML 2024 · 183 citations
- Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsYongqi An, Xu Zhao, Tao Yu, Ming Tang et al.AAAI 2024 · 130 citations
Related papers
- Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-ProbingQi Le, Enmao Diao, Ziyan Wang, Xinran Wang et al.ICLR 2025
- DLP: Dynamic Layerwise Pruning in Large Language ModelsYuli Chen, Bo Cheng, Jiale Han, Yingying Zhang et al.ICML 2025
- Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution OptimizationGuanchen Li, Yixing Xu, Zeping Li, Ji Liu et al.NeurIPS 2025 · 7 citations
- Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMsChang Gao, Kang Zhao, Runqi Wang, Jianfei Chen et al.ACM MM 2025
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language ModelsZiyan Wang, Enmao Diao, Qi Le, Pu Wang et al.ACL 2026 · 2 citations
