SC2025Top-tier venue
KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPU
Hemeng Wang, Yang Du, Sidu Li, Xiaowen Tian, Qingxiao Sun, Weifeng Liu
Abstract
Efficient general matrix-matrix multiplication (GEMM) has attracted significant research attention in HPC and AI workloads. While large-scale GEMM has nearly achieved the peak floatingpoint performance of GPUs, substantial opportunities for optimization remain in small and batched GEMM operations.
We in this paper propose KAMI, a set of 1D, 2D, and 3D GEMM algorithms that extend the theory of communication-avoiding (CA) techniques within a single GPU. KAMI optimizes thread blocklevel GEMM by utilizing tensor cores as computational units, lowlatency thread registers as local memory, and high-latency onchip shared memory as a communication medium. We provide a theoretical analysis of CA performance from the perspective of GPU clock cycles, rather than the traditional execution time. Also, we implement sparse-dense matrix-matrix multiplication (SpMM) and sparse general matrix-matrix multiplication (SpGEMM) with this compute-communication pattern. Experimental results for general, low-rank, batched, and sparse multiplication on NVIDIA, AMD, and Intel GPUs show significant performance improvements over existing libraries cuBLAS, cuBLASDx, CUTLASS, MAGMA, and SYCL-Bench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8b8465e-8dc3-480b-95cf-f3ef74b174a9Builds on27
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
- GE-SpMM: general-purpose sparse matrix-matrix multiplication on GPUs for graph neural networksGuyue Huang, Guohao Dai, Yu Wang, Huazhong YangSC 2020 · 130 citations
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 67 citations
- TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUsYuyao Niu, Zhengyang Lu, Haonan Ji, Shuhui Song et al.PPoPP 2022 · 66 citations
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu et al.SC 2020 · 65 citations
Related papers
- CA3DMM: A New Algorithm Based on a Unified View of Parallel Matrix MultiplicationHua Huang, Edmond ChowSC 2022 · 4 citations
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- HSMU-SpGEMM: Achieving High Shared Memory Utilization for Parallel Sparse General Matrix-Matrix Multiplication on Modern GPUsMin Wu, Huizhang Luo, Fenfang Li, Yiran Zhang et al.HPCA 2025 · 3 citations
- VSpGEMM: Exploiting Versal ACAP for High-Performance SpGEMM AccelerationKai Shi, Zhe Lin, Xinya Luan, Jianwang Zhai et al.DAC 2025 · 1 citation
- HyTiS: Hybrid Tile Scheduling for GPU GEMM with Enhanced Wave Utilization and Cache LocalityZheng Zhang, Hulin Wang, Hongming Xu, Donglin Yang et al.SC 2025 · 4 citations
