Lune

NeurIPS2025顶会

Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity

Susav Shrestha, Bradley W. Settlemyer, Nikoli Dryden, A. L. Narasimha Reddy

2025年份
8被引次数
1顶会引用

摘要

Accelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamically activates only a small subset of the model parameters, shows promise but does not scale to large batch sizes due to union of active neurons quickly approaching dense computation. We introduce Polar Sparsity, highlighting a key shift in sparsity importance from MLP to Attention layers as we scale batch size and sequence length. While MLP layers become more compute-efficient under batching, their sparsity vanishes. In contrast, attention becomes increasingly more expensive at scale, while their head sparsity remains stable and batch-invariant. We develop Selective Head Attention with hardware-efficient, sparsity-aware GPU kernels, delivering up to 2.2× end-to-end speedups for models like OPT, LLaMA-2 & 3, Qwen, Mistral across various batch sizes and sequence lengths without compromising accuracy. To our knowledge, this is the first work to demonstrate that contextual sparsity can scale effectively to large batch sizes, delivering substantial inference acceleration with minimal changes, making Polar Sparsity practical for large-scale, high-throughput LLM deployment systems. Our code is available at: https://github.com/susavlsh10/Polar-Sparsity.

We leverage these properties to build contextual sparsity-aware Selective GEMM and Selective FlashAttention kernels that reduce memory I/O and compute, enabling scalable and high-throughput inference. To the best of our knowledge, we are the first work to show that contextual sparsity is scalable with batch size, offering even higher gains at larger batch sizes. Our main contributions are as follows:

  1. We show that activation sparsity in the MLP layers degrades with batch size due to union activations, while attention head sparsity remains stable and batch-invariant.

  2. We design Selective GEMM kernels with a layer-wise top-k optimization strategy for dynamic MLP activations, achieving up to 5.5× speedup.

  3. We introduce Selective FlashAttention kernels that support Head/Group sparsity with perquery activation, reducing memory I/O and compute, achieving up to 2.8× speedup.

  4. Polar Sparsity delivers up to 2.2× improvement in batched decoding throughput with negligible accuracy loss, and can be seamlessly integrated into a wide range of LLMs, including those without ReLU activations, to unlock substantial inference acceleration.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 161ca0e5-26ec-414d-be4e-ee31b927cd8a

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖