Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari
摘要
In the era of large language models (LLMs) and long-context generation, model compression techniques such as pruning, quantization, and distillation offer effective ways to reduce memory usage.Among them, pruning is constrained by the difficulty of exploiting unstructured sparsity on modern hardware.Consequently, LLM pruning is often restricted to structured patterns for hardware efficiency, although unstructured sparsity offers better accuracy retention at higher sparsity.To bridge this gap between the full potential of pruning and efficiency, we propose the Coruscant GPU SpMM kernel that leverages a bitmap-based sparse format for reduced memory footprint inside GPU memory and reduced latency of memory-bound matrix multiplications in LLM inference.This is achieved by transferring the compressed matrix tiles to GPU processors and decompressing them locally for tensor core execution.We see further optimization opportunity in microarchitecture-level and propose Coruscant Sparse Tensor Core, which computes directly on the compressed format without decompression by integrating a bitmap decoder.Coruscant kernel achieves up to 2× speedup over cuBLAS and 1.48× over Flash-LLM.With Coruscant Sparse Tensor Core, the speedup reaches 2.75× over cuBLAS.Most importantly, Coruscant serves as an ideal solution for state-of-the-art LLM pruning methods by significantly reducing the memory footprint and accelerating SpMM on sparsity range 30% to 70%, enabling exploration of diverse sparsity patterns and pruning strategies.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 被引用 12 次
- Early Silicon of Raptor: The First 3D-DRAM Accelerator for Generative InferencePrashant J. Nair, Ramyad Hadidi, Subramani Ganesh, Sangamesh Kodge 等ISCA 2026 · 被引用 4 次
- Unified Static-Dynamic Pruning for Efficient LLM InferenceJinhyeok Kim, Yejoon Lee, Jaeyoung DoVLDB 2026
相关 Paper
- SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsRuibo Fan, Xiangrui Yu, Peijie Dong, Zeyu Li 等EuroSys 2025 · 被引用 22 次
- Flash-LLM: Enabling Low-Cost and Highly-Efficient Large Generative Model Inference With Unstructured SparsityHaojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang 等VLDB 2024 · 被引用 29 次
- DELTA4: Sparse Matrix-Vector Multiplication for Low SparsityVladimír Macko, Vladimír BožaICML 2026 · 被引用 9 次
- Taming Unstructured Sparsity on GPUs via Latency-Aware OptimizationMaohua Zhu, Yuan XieDAC 2020 · 被引用 3 次
- GeneralSparse: Bridging the Gap in SpMM for Pruned Large Language Model Inference on GPUsYaoyu Wang, Xiao Guo, Junmin Xiao, De Chen 等USENIX ATC 2025 · 被引用 5 次
