Lune

NeurIPS2025Top-tier venue

Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity

Susav Shrestha, Bradley W. Settlemyer, Nikoli Dryden, A. L. Narasimha Reddy

2025Year
8Citations
1Top-tier citations

Abstract

Accelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamically activates only a small subset of the model parameters, shows promise but does not scale to large batch sizes due to union of active neurons quickly approaching dense computation. We introduce Polar Sparsity, highlighting a key shift in sparsity importance from MLP to Attention layers as we scale batch size and sequence length. While MLP layers become more compute-efficient under batching, their sparsity vanishes. In contrast, attention becomes increasingly more expensive at scale, while their head sparsity remains stable and batch-invariant. We develop Selective Head Attention with hardware-efficient, sparsity-aware GPU kernels, delivering up to 2.2× end-to-end speedups for models like OPT, LLaMA-2 & 3, Qwen, Mistral across various batch sizes and sequence lengths without compromising accuracy. To our knowledge, this is the first work to demonstrate that contextual sparsity can scale effectively to large batch sizes, delivering substantial inference acceleration with minimal changes, making Polar Sparsity practical for large-scale, high-throughput LLM deployment systems. Our code is available at: https://github.com/susavlsh10/Polar-Sparsity.

We leverage these properties to build contextual sparsity-aware Selective GEMM and Selective FlashAttention kernels that reduce memory I/O and compute, enabling scalable and high-throughput inference. To the best of our knowledge, we are the first work to show that contextual sparsity is scalable with batch size, offering even higher gains at larger batch sizes. Our main contributions are as follows:

  1. We show that activation sparsity in the MLP layers degrades with batch size due to union activations, while attention head sparsity remains stable and batch-invariant.

  2. We design Selective GEMM kernels with a layer-wise top-k optimization strategy for dynamic MLP activations, achieving up to 5.5× speedup.

  3. We introduce Selective FlashAttention kernels that support Head/Group sparsity with perquery activation, reducing memory I/O and compute, achieving up to 2.8× speedup.

  4. Polar Sparsity delivers up to 2.2× improvement in batched decoding throughput with negligible accuracy loss, and can be seamlessly integrated into a wide range of LLMs, including those without ReLU activations, to unlock substantial inference acceleration.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 161ca0e5-26ec-414d-be4e-ee31b927cd8a

Cited by top-tier papers1

Ask how each one uses it

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines