Eureka: Efficient Tensor Cores for One-sided Unstructured Sparsity in DNN Inference
Ashish Gondimalla, Mithuna Thottethodi, T. N. Vijaykumar
Abstract
Deep neural networks (DNNs), while enormously popular, continue to place ever higher compute demand for which GPUs provide specialized matrix multipliers called tensor cores. To reduce the compute demand via sparsity, Nvidia Ampere’s tensor cores support 2:4 structured sparsity in the filters (i.e., two non-zeros out of four values) which provides uniform 50% sparsity without any load imbalance issues. Consequently, the sparse tensor cores maintain (input or output) operand stationarity, which is fundamental for avoiding high-overhead hardware, requiring only one extra 4-1 multiplexer per multiply-accumulate unit (MAC). However, 2:4 sparsity is limited to 2x improvements in performance and energy without loss of accuracy, whereas unstructured sparsity provides 5-6x opportunity albeit while causing load imbalance. Previous papers on unstructured sparsity incur high hardware overhead (e.g., buffering, crossbars, scatter-gather networks, and address calculators) mainly due to sacrificing operand stationarity in favor of load balance. To avoid adding high overheads to the highly-efficient tensor cores, we propose Eureka, an efficient tensor core for unstructured sparsity. Eureka addresses load imbalance via three contributions: (1) Our key insight is that a slight weakening of output stationarity achieves load balance most of the time while incurring only a modest hardware overhead. Accordingly, we propose single-step uni-directional displacement (SUDS), where a filter element’s multiplication can either occur in its original position or be displaced to a vacant MAC in the adjacent row below while the accumulation occurs in the original row to restore output stationarity. SUDS is an offline technique for inference. (2) We provide an optimal algorithm for work assignment for SUDS. (3) To achieve fewer bubbles in the tensor core’s systolic pipeline due to the irregularity of unstructured sparsity, we propose offline systolic scheduling to group together the sparse filters with similar, statically-known execution times (based on the number of non-zeros). Our evaluation shows that Eureka achieves 4.8x and 2.4x speedups, and 3.1x and 1.8x energy reductions over dense and 2:4 sparse (Ampere) implementations, respectively, and incurs area and power overheads of 6% and 11.5%, respectively, over Ampere.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c5d81d4b-c5ae-41c0-89fa-a64abb00a07cCited by top-tier papers7
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue et al.MICRO 2024 · 31 citations
- BBS: Bi-Directional Bit-Level Sparsity for Deep Learning AccelerationYuzong Chen, Jian Meng, Jae-sun Seo, Mohamed S. AbdelfattahMICRO 2024 · 25 citations
- BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationYuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang et al.HPCA 2025 · 23 citations
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng et al.ISCA 2025 · 17 citations
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park et al.HPCA 2025 · 9 citations
Related papers
- Taming Unstructured Sparsity on GPUs via Latency-Aware OptimizationMaohua Zhu, Yuan XieDAC 2020 · 3 citations
- Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruningGuyue Huang, Haoran Li, Minghai Qin, Fei Sun et al.DAC 2022 · 19 citations
- RM-STC: Row-Merge Dataflow Inspired GPU Sparse Tensor Core for Energy-Efficient Sparse AccelerationGuyue Huang, Zhengyang Wang, Po-An Tsai, Chen Zhang et al.MICRO 2023 · 15 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta et al.SC 2024 · 15 citations
