Lune

ICML2026顶会

DELTA4: Sparse Matrix-Vector Multiplication for Low Sparsity

Vladimír Macko, Vladimír Boža

2026年份
9被引次数

摘要

Sparse Matrix-Vector Multiplication (SpMV) is a fundamental operation in the inference of sparse Large Language Models (LLMs). Because existing SpMV methods perform poorly under the low, unstructured sparsity (30−9030-90\\%) commonly observed in pruned LLMs, unstructured pruning provides only limited memory reduction and speedup. We propose DELTA4-SpMV, a GPU-optimized format and kernel co-designed to reduce storage overhead while remaining compatible with the GPU’s execution model. This enables efficient SpMV for unstructured sparsity without specialized hardware units or precomputation. We identify memory bandwidth as the primary limiting factor of SpMV and analyze the storage overhead of DELTA4. At 5050\\% sparsity, DELTA4 is the first approach to achieve 1.5×1.5\times memory reduction and 1.2−1.5×1.2-1.5\times speedup over the dense baseline as well as substantial improvements over other SpMV methods: cuSPARSE (2.8−13.0×2.8-13.0\times), Sputnik (1.9−2.6×1.9-2.6\times), and DASP (2.2−2.5×2.2-2.5\times). An LLM pruned with Wanda to sparsity 5050\\% requires 1.5×1.5\times less memory and achieves 1.5×1.5\times faster inference at fp16 precision. As a result, unstructured pruning at 5050\\% sparsity becomes practical for real-world LLM workloads and bridges the efficiency gap with structured 2:4 sparsity.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖