Lune

MICRO2025顶会

Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference

Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari

2025年份
8被引次数
3顶会引用

摘要

In the era of large language models (LLMs) and long-context generation, model compression techniques such as pruning, quantization, and distillation offer effective ways to reduce memory usage.Among them, pruning is constrained by the difficulty of exploiting unstructured sparsity on modern hardware.Consequently, LLM pruning is often restricted to structured patterns for hardware efficiency, although unstructured sparsity offers better accuracy retention at higher sparsity.To bridge this gap between the full potential of pruning and efficiency, we propose the Coruscant GPU SpMM kernel that leverages a bitmap-based sparse format for reduced memory footprint inside GPU memory and reduced latency of memory-bound matrix multiplications in LLM inference.This is achieved by transferring the compressed matrix tiles to GPU processors and decompressing them locally for tensor core execution.We see further optimization opportunity in microarchitecture-level and propose Coruscant Sparse Tensor Core, which computes directly on the compressed format without decompression by integrating a bitmap decoder.Coruscant kernel achieves up to 2× speedup over cuBLAS and 1.48× over Flash-LLM.With Coruscant Sparse Tensor Core, the speedup reaches 2.75× over cuBLAS.Most importantly, Coruscant serves as an ideal solution for state-of-the-art LLM pruning methods by significantly reducing the memory footprint and accelerating SpMM on sparsity range 30% to 70%, enabling exploration of diverse sparsity patterns and pruning strategies.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 64b1044d-43e2-48e7-971f-37d4e08f2beb

引用它的顶会 Paper3

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖