SlideSparse: Fast and Flexible (2N-2):2N Structured Sparsity
Yingbo HAO, Hanyong Shao, Ting Song, Yan Xia, Di Zhang, Shaohan Huang, Xun Wu, Songchen Xu, Le Xu, Li Dong, Zewen Chi, Yi Zou, Furu Wei
Abstract
NVIDIA's 2:4 Sparse Tensor Cores deliver throughput but demand strict 50% pruning—a ratio that causes severe accuracy loss in LLMs. Milder patterns (e.g., 6:8, 25% pruning) preserve accuracy far better—within 0.4–1.8 average points of dense in our Qwen2.5-7B/14B study—yet receive NO hardware support and fall back to dense execution. We present SlideSparse, the first system to unlock Sparse Tensor Core acceleration for the model family on commodity GPUs. Our Sliding Window Decomposition rewrites any weight block into overlapping 2:4-compliant windows without changing the underlying dot product; in addition, our Activation Lifting fuses the corresponding activation rearrangement into per-token quantization at low marginal cost. Integrated into vLLM, SlideSparse is evaluated across various GPUs (A100, H100, B200, RTX 4090, RTX 5080, DGX-spark), precisions (FP4, INT8, FP8, BF16, FP16), and model families (Llama, Qwen, BitNet). On compute-bound workloads, the measured speedup () matches the theoretical upper-bound at 6:8 weight sparsity in Qwen2.5-7B, establishing as a practical path to better accuracy–speedup trade-offs in LLM acceleration. Code available at https://github.com/bcacdwk/vllmbench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 412119f4-ea04-4467-8cbb-4c3fff249708Builds on7
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
Related papers
- Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruningGuyue Huang, Haoran Li, Minghai Qin, Fei Sun et al.DAC 2022 · 19 citations
- TSENOR: Highly-Efficient Algorithm for Finding Transposable N: M Sparse MasksXiang Meng, Mehdi Makni, Rahul MazumderNeurIPS 2025
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language ModelsWenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang et al.ICLR 2026 · 12 citations
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu et al.ASPLOS 2025 · 5 citations
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix FactorizationLawrence Liu, Alexander Liu, Mengdi Wang, Tuo Zhao et al.ICLR 2026 · 3 citations
