HotTiles: Accelerating SpMM with Heterogeneous Accelerator Architectures
Gerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim Hur
Abstract
Sparse Matrix Dense Matrix Multiplication (SpMM) is an important kernel with application across a wide range of domains, including machine learning and linear algebra solvers. In many sparse matrices, the pattern of nonzeros is nonuniform: nonzeros form dense and sparse regions, rather than being uniformly distributed across the whole matrix. We refer to this property as Intra-Matrix Heterogeneity (IMH). Currently, SpMM accelerator designs do not leverage this heterogeneity. They employ the same processing elements (PEs) for all the regions of a sparse matrix, resulting in suboptimal acceleration. To address this limitation, we utilize heterogeneous SpMM accelerator architectures, which include different types of PEs to exploit IMH. We develop an analytical modeling framework to predict the performance of different types of accelerator PEs taking into account IMH. Furthermore, we present a heuristic for partitioning sparse matrices among heterogeneous PEs. We call our matrix modeling and partitioning method HotTiles. To evaluate HotTiles, we simulate three different heterogeneous architectures. Each one consists of two types of workers (i.e., PEs): one suited for compute-bound denser regions (Hot Worker) and one for memory-bound sparser regions (Cold Worker). Our results show that exploiting IMH with HotTiles is very effective. Depending on the architecture, heterogeneous execution with HotTiles outperforms homogeneous execution using only hot or only cold workers by 9.2-16.8x and 1.4-3.7x, respectively. In addition, HotTiles outperforms the best worker type used on a per-matrix basis by 1.3-2.5 x. Finally, HotTiles outperforms an IMH-unaware heterogeneous execution strategy by 1.4-2.2x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cc42204-1880-4e8a-8f24-0bbe8d13b8eeCited by top-tier papers6
- HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU CoresZhonggen Li, Xiangyu Ke, Yifan Zhu, Yunjun Gao et al.ICDE 2025 · 5 citations
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman et al.MICRO 2025 · 5 citations
- Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation FrameworkAmir Ghazizadeh Ahsaei, Lingxiang Yin, Shilin Tian, Fangzhou Ye et al.MICRO 2025 · 4 citations
- HYTE: Flexible Tiling for Sparse Accelerators via Hybrid Static-Dynamic ApproachesXintong Li, Zhiyao Li, Mingyu GaoISCA 2025 · 2 citations
- NetSparse: In-Network Acceleration of Distributed Sparse KernelsGerasimos Gerogiannis, Dimitrios Merkouriadis, Charles Block, Annus Zulfiqar et al.MICRO 2025 · 1 citation
Builds on16
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen et al.ASPLOS 2020 · 171 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna et al.HPCA 2021 · 143 citations
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong et al.HPCA 2020 · 121 citations
Related papers
- ACES: Accelerating Sparse Matrix Multiplication with Adaptive Execution Flow and Concurrency-Aware Cache OptimizationsXiaoyang Lu, Boyu Long, Xiaoming Chen, Yinhe Han et al.ASPLOS 2024 · 13 citations
- SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMMGerasimos Gerogiannis, Serif Yesil, Damitha Lenadora, Dingyuan Cao et al.ISCA 2023 · 27 citations
- SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationHaoqin Huang, Pengcheng Yao, Zhaozeng An, Yufei Sun et al.DAC 2024 · 3 citations
- ASM-SpMM: Unleashing the Potential of Arm SME for Sparse Matrix Multiplication AccelerationJiazhi Jiang, Xijia Yao, Jiayu Chen, Jinhui Wei et al.PPoPP 2026
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta et al.SC 2024 · 15 citations
