HiT: A Unified Sparsity-Adaptive Architecture for High-Throughput Matrix Multiplication
Tingting Xiang, Xiaochen Wang, Miao Yu, Trevor E. Carlson
Abstract
Accelerating matrix operations has become increasingly critical as AI models and scientific workloads continue to scale. These applications involve matrices spanning sparsity levels from <0.0001% to fully dense, demanding accelerators that maintain high performance across this full range. However, prior designs either target a narrow sparsity range, resulting in inefficiencies outside their specialization, or support broad sparsity at the cost of throughput, with the state-of-the-art accelerator achieving less than 3.125% of peak performance on highly sparse matrices. We present HiT, a unified sparsity-adaptive architecture that delivers consistently high throughput and efficiency across the entire sparsity spectrum. HiT introduces two novel outer-productbased dataflows, HSparse and MSparse, supported by a Parallel Intersection & Distribution Unit and a Dual-mode Accumulator, targeting highly and moderately sparse workloads, respectively. These dataflows enable regular memory access to sparse data and on-chip accumulation of partial sums while exploiting two levels of spatial parallelism. As a result, HiT achieves high intersection rates (more valid non-zero matches per cycle) and data reuse, leading to high throughput. For dense workloads, HiT employs an inner-product dataflow to maximize compute efficiency. We benchmark HiT against Trapezoid, a state-of-the-art accelerator for full-spectrum sparsity. Specifically, it delivers a geomean performance/area improvement on highly sparse × highly sparse multiplications, 2.18× across all highly sparse workloads, and on moderately sparse workloads. Across a comprehensive set of dense and sparse workloads, HiT achieves higher geomean performance/area and reduces energy consumption by compared to Trapezoid.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4346ede6-bb52-4613-a678-cc6903e374aeBuilds on24
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingTong Geng, Ang Li, Runbin Shi, Chunshu Wu et al.MICRO 2020 · 299 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
Related papers
- Trapezoid: A Versatile Accelerator for Dense and Sparse Matrix MultiplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2024 · 38 citations
- Harmonia: A Unified Hierarchical Scheduling Framework for Sparse Matrix MultiplicationJingkui Yang, Fangxin Liu, Xin Ju, Ning Yang et al.ISCA 2026
- DenSparSA: A Balanced Systolic Array Approach for Dense and Sparse Matrix MultiplicationZiheng Wang, Ruiqi Sun, Xin He, Tianrui Ma et al.DAC 2025 · 1 citation
- SLAWS: Spatial Locality Analysis and Workload Orchestration for Sparse Matrix MultiplicationGuoyu Li, Zheng Guan, Beichen Zhang, Jun Yu et al.ASPLOS 2026
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi et al.MICRO 2020 · 223 citations
