Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
Yuqi Luo, Chenyang Song, Xu Han, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu, Maosong Sun
Abstract
Activation sparsity denotes the existence of substantial weakly-contributed neurons within feedforward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1× speedup compared with its dense version. The codes and checkpoints are available at https: //github.com/thunlp/SparsingLaw/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b9a45a3-0daa-4a46-9af5-4cd045579d0fCited by top-tier papers6
- Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual SparsitySusav Shrestha, Bradley W. Settlemyer, Nikoli Dryden, A. L. Narasimha ReddyNeurIPS 2025 · 8 citations
- Universal Properties of Activation Sparsity in Modern Large Language ModelsFilip Szatkowski, Patryk Będkowski, Alessio Devoto, Jan Dubiński et al.ICLR 2026 · 5 citations
- Spark Transformer: Reactivating Sparsity in Transformer FFN and AttentionChong You, Kan Wu, Zhipeng Jia, Lin Chen et al.NeurIPS 2025 · 3 citations
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph EmbeddingYifei Li, Lingling Zhang, Hang Yan, Tianzhe Zhao et al.KDD 2025 · 2 citations
- Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation LinkingTuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao et al.ASPLOS 2025 · 1 citation
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou et al.ICML 2023 · 318 citations
- Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural NetworksMark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev et al.ICML 2020 · 163 citations
Related papers
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao et al.NeurIPS 2024 · 29 citations
- Training-Free Activation Sparsity in Large Language ModelsJames Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo et al.ICLR 2025
- Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language ModelsMing Wang, Miao Zhang, Xuebo Liu, Liqiang NieEMNLP 2025
- SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model CompressionXinhao Huang, You-Liang Huang, Zeyi WenAAAI 2025 · 14 citations
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM InferenceZhenyu Zhang, Zechun Liu, Yuandong Tian, Harshit Khaitan et al.ICLR 2025
