Grasp: Group-based Prediction of Activation Sparsity for Fast LLM Inference
Jiho Shin, Hoeseok Yang, Youngmin Yi
Abstract
Optimizing LLM inference has become increasingly important as the demand for efficient on-device deployments grows. To reduce the computational overhead in the MLP components, which account for a significant portion of LLM inference, ReLU-fied LLMs have been introduced to maximize activation sparsity. Several sparsity prediction methods have been developed to efficiently skip unnecessary memory accesses and computations by predicting activation sparsity. In this paper, we propose a novel magnitude-based, training-free sparsity prediction technique called Grasp that builds on the existing sign bitbased method for ReLU-fied LLMs. The proposed method enhances prediction accuracy by grouping values considering the distribution within vectors and explicitly accounting for statistical outliers. This allows us to estimate the impact of each element more accurately yet in an efficient way, improving both activation sparsity prediction accuracy and computational efficiency. Compared to the-state-of-the-art technique, Grasp achieves higher sparsity prediction accuracy and higher skipping efficiency, which corresponds to speedup against the dense inference.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- oFFN: Outlier and Neuron-aware Structured FFN for Fast yet Accurate LLM InferenceGeunsoo Song, Hoeseok Yang, Youngmin YiASPLOS 2026
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM InferenceZhenyu Zhang, Zechun Liu, Yuandong Tian, Harshit Khaitan et al.ICLR 2025
- Training-Free Activation Sparsity in Large Language ModelsJames Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo et al.ICLR 2025
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model InferenceSihan Chen, Dan Zhao, Jongwoo Ko, Colby Banbury et al.ICLR 2026 · 3 citations
- Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language ModelsMing Wang, Miao Zhang, Xuebo Liu, Liqiang NieEMNLP 2025
