RM-STC: Row-Merge Dataflow Inspired GPU Sparse Tensor Core for Energy-Efficient Sparse Acceleration
Guyue Huang, Zhengyang Wang, Po-An Tsai, Chen Zhang, Yufei Ding, Yuan Xie
Abstract
This paper proposes RM-STC, a novel GPU tensor core architecture designed for sparse Deep Neural Networks (DNNs) with two key innovations: (1) native support for both training and inference and (2) high efficiency for all sparsity degrees. To achieve the first goal, RM-STC employs a uniform sparse encoding scheme that natively supports all operations holistically in forward and backward passes, thereby eliminating the need for costly sparse encoding transformation in between. For the second goal, RM-STC takes inspiration from the row-merge dataflow and combines the input-gathering and output-scattering hardware features to minimize the energy overhead. Experiments show that RM-STC achieves significant speedups and energy efficiency improvements over dense tensor cores and previous sparse tensor cores.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue et al.MICRO 2024 · 31 citations
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng et al.ISCA 2025 · 17 citations
- TB-STC: Transposable Block-wise N: M Structured Sparse Tensor CoreJun Liu, Shulin Zeng, Junbo Zhao, Li Ding et al.HPCA 2025 · 9 citations
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model ServingJungi Lee, Junyong Park, Soohyun Cha, Jaehoon Cho et al.MICRO 2025 · 7 citations
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionHuizheng Wang, Hongbin Wang, Zichuan Wang, Zhiheng Yue et al.HPCA 2026 · 2 citations
Related papers
- Uni-STC: Unified Sparse Tensor CoreHaocheng Lian, Qiyue Zhang, Xinran Zhao, Meichen Dong et al.HPCA 2026 · 2 citations
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo et al.ISCA 2021 · 109 citations
- Taming Unstructured Sparsity on GPUs via Latency-Aware OptimizationMaohua Zhu, Yuan XieDAC 2020 · 3 citations
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang et al.DAC 2020 · 12 citations
- Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruningGuyue Huang, Haoran Li, Minghai Qin, Fei Sun et al.DAC 2022 · 19 citations
