Dual Activation-Weight Sparsity: A Training-Free Framework for Efficient Large Language Model Compression
Luoyang Sun, Guangyan Li, Cheng Deng, Haifeng Zhang, Jian Zhao, Yongqiang Tang, Wensheng Zhang, Jun Wang
Abstract
Large language models (LLMs) excel at natural language tasks but face deployment challenges due to computational demands. We introduce Dual Activation-Weight Sparsity (DAWS), a training-free framework that jointly exploits activation and weight sparsity through magnitude-based routing. Systematic analysis of pretrained transformers reveals two key observations (Figure 1 ): (1) the activation energy is concentrated in a few neurons, and (2) activation-weight contribution patterns exhibit substantial module-wise heterogeneity across attention and FFN modules. DAWS employs a three-tier routing strategy: high-magnitude activations pass through fullprecision weights to preserve critical pathways, medium-magnitude activations use magnitudepruned sparse weights for efficiency, and lowmagnitude activations are directly discarded. Unlike prior work that uses activation-aware pruning methods like WANDA, our approach uses direct magnitude-based pruning, which we show is more robust to sample-level variations. Experiments on Llama and Mistral models demonstrate that DAWS maintains >98% of dense model performance at 50% sparsity, outperforming WANDA, TEAL, and R-Sparse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85dd457a-6b4e-4a1b-8900-744115bcf6b2Builds on9
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
Related papers
- Training-Free Activation Sparsity in Large Language ModelsJames Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo et al.ICLR 2025
- SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language ModelsXun Liang, Hanyu Wang, Huayi Lai, Simin Niu et al.AAAI 2026
- Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language ModelsMing Wang, Miao Zhang, Xuebo Liu, Liqiang NieEMNLP 2025
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model InferenceSihan Chen, Dan Zhao, Jongwoo Ko, Colby Banbury et al.ICLR 2026 · 3 citations
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM InferenceZhenyu Zhang, Zechun Liu, Yuandong Tian, Harshit Khaitan et al.ICLR 2025
