ICML2026
SlideSparse: Fast and Flexible (2N-2):2N Structured Sparsity
Yingbo HAO, Hanyong Shao, Ting Song, Yan Xia, Di Zhang, Shaohan Huang, Xun Wu, Songchen Xu, Le Xu, Li Dong, Zewen Chi, Yi Zou, Furu Wei
Abstract
NVIDIA's 2:4 Sparse Tensor Cores deliver throughput but demand strict 50% pruning—a ratio that causes severe accuracy loss in LLMs. Milder patterns (e.g., 6:8, 25% pruning) preserve accuracy far better—within 0.4–1.8 average points of dense in our Qwen2.5-7B/14B study—yet receive NO hardware support and fall back to dense execution. We present SlideSparse, the first system to unlock Sparse Tensor Core acceleration for the model family on commodity GPUs. Our Sliding Window Decomposition rewrites any weight block into overlapping 2:4-compliant windows without changing the underlying dot product; in addition, our Activation Lifting fuses the corresponding activation rearrangement into per-token quantization at low marginal cost. Integrated into vLLM, SlideSparse is evaluated across various GPUs (A100, H100, B200, RTX 4090, RTX 5080, DGX-spark), precisions (FP4, INT8, FP8, BF16, FP16), and model families (Llama, Qwen, BitNet). On compute-bound workloads, the measured speedup () matches the theoretical upper-bound at 6:8 weight sparsity in Qwen2.5-7B, establishing as a practical path to better accuracy–speedup trade-offs in LLM acceleration. Code available at https://github.com/bcacdwk/vllmbench.