SparseInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
Qinsi Wang, Saeed Vahidian, Hancheng Ye, Jianyang Gu, Jianyi Zhang, Yiran Chen
Abstract
Large language models (LLMs) with billions of parameters have sparked a new wave of exciting AI applications. However, their high computational costs and memory demands during inference pose significant challenges. Adaptive sparse activation inference, which activates only a small number of neurons for each token, offers a novel way to accelerate model inference without degrading performance, showing great potential for resource-constrained hardware devices. Nevertheless, existing methods predict activated neurons based on individual tokens with additional MLP, which involve frequent changes in activation maps and resource calls, limiting the acceleration benefits of sparse activation. In this paper, we introduce CoreInfer, an MLP-free adaptive sparse activation inference method based on sentence-level prediction. Specifically, we propose the concept of sentence-wise core neurons, which refers to the subset of neurons most critical for a given sentence, and empirically demonstrate its effectiveness. To determine the core neurons, we explore the correlation between core neurons and the sentence's semantics. Remarkably, we discovered that core neurons exhibit both stability and similarity in relation to the sentence's semantics-an insight overlooked by previous studies. Building on this finding, we further design two semantic-based methods for predicting core neurons to fit different input scenarios. In CoreInfer, the core neurons are determined during the pre-filling stage and fixed during the encoding stage, enabling zero-cost sparse inference. We evaluated the model generalization and task generalization of CoreInfer across various models and tasks. Notably, on an NVIDIA TITAN XP GPU, CoreInfer achieved a 10.33×and 2.72×speedup compared to the Huggingface implementation and PowerInfer, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 874d95fa-f221-4710-8fbe-e38207da4e2fCited by top-tier papers4
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent SystemsHancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang et al.NeurIPS 2025 · 42 citations
- Angles Don't Lie: Unlocking Training‑Efficient RL Through the Model's Own SignalsQinsi Wang, Jinghan Ke, Hancheng Ye, Yueqian Lin et al.NeurIPS 2025 · 16 citations
- CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language ModelsQinsi Wang, Hancheng Ye, Ming-Yu Chung, Yudong Liu et al.ICML 2025
- Seeing is Solving: Unlocking Efficient Multimodal RL via View AlignmentQinsi Wang, Jing Shi, Kun Wan, Handong Zhao et al.ICML 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou et al.ICML 2023 · 318 citations
Related papers
- PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPUYixin Song, Zeyu Mi, Haotong Xie, Haibo ChenSOSP 2024 · 86 citations
- DynamicInfer: Runtime-Aware Sparse Offloading for LLMs Inference on a Consumer-Grade GPUZhui Zhu, Weichen Zhang, Zhenghan Zhou, Yunhao Liu et al.ICLR 2026
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model InferenceSihan Chen, Dan Zhao, Jongwoo Ko, Colby Banbury et al.ICLR 2026 · 3 citations
- COUNTDOWN: Contextually Sparse Activation Filtering Out Unnecessary Weights in Down ProjectionJaewon Cheon, Pilsung KangEMNLP 2025
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language ModelsIman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C. del Mundo et al.ICLR 2024 · 109 citations
