DELTA4: Sparse Matrix-Vector Multiplication for Low Sparsity
Vladimír Macko, Vladimír Boža
摘要
Sparse Matrix-Vector Multiplication (SpMV) is a fundamental operation in the inference of sparse Large Language Models (LLMs). Because existing SpMV methods perform poorly under the low, unstructured sparsity () commonly observed in pruned LLMs, unstructured pruning provides only limited memory reduction and speedup. We propose DELTA4-SpMV, a GPU-optimized format and kernel co-designed to reduce storage overhead while remaining compatible with the GPU’s execution model. This enables efficient SpMV for unstructured sparsity without specialized hardware units or precomputation. We identify memory bandwidth as the primary limiting factor of SpMV and analyze the storage overhead of DELTA4. At sparsity, DELTA4 is the first approach to achieve memory reduction and speedup over the dense baseline as well as substantial improvements over other SpMV methods: cuSPARSE (), Sputnik (), and DASP (). An LLM pruned with Wanda to sparsity requires less memory and achieves faster inference at fp16 precision. As a result, unstructured pruning at sparsity becomes practical for real-world LLM workloads and bridges the efficiency gap with structured 2:4 sparsity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li 等ICML 2023 · 被引用 683 次
相关 Paper
- SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsRuibo Fan, Xiangrui Yu, Peijie Dong, Zeyu Li 等EuroSys 2025 · 被引用 22 次
- Flash-LLM: Enabling Low-Cost and Highly-Efficient Large Generative Model Inference With Unstructured SparsityHaojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang 等VLDB 2024 · 被引用 29 次
- GeneralSparse: Bridging the Gap in SpMM for Pruned Large Language Model Inference on GPUsYaoyu Wang, Xiao Guo, Junmin Xiao, De Chen 等USENIX ATC 2025 · 被引用 5 次
- Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariMICRO 2025 · 被引用 8 次
- Unified Static-Dynamic Pruning for Efficient LLM InferenceJinhyeok Kim, Yejoon Lee, Jaeyoung DoVLDB 2026
