SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMM
Yunan Zhang, Po-An Tsai, Hung-Wei Tseng
Abstract
Matrix-multiplication units (MXUs) are now prevalent in every computing platform. The key attribute that makes MXUs so successful is the semiring structure, which allows tiling for both parallelism and data reuse. Nonetheless, matrix-multiplication is not the only algorithm with such attributes. We find that many algorithms share the same structure and differ in only the core operation; for example, using add-minimum instead of multiply-add. Algorithms with a semiring-like structure therefore have potential to be accelerated by a general-purpose matrix operation architecture, instead of common MXUs.
In this paper, we propose SIMD 2 , a new programming paradigm to support generalized matrix operations with a semiring-like structure. SIMD 2 instructions accelerate eight more types of matrix operations, in addition to matrix multiplications. Since SIMD 2 instructions resemble a matrix-multiplication instruction, we are able to build SIMD 2 architecture on top of any MXU architecture with minimal modifications. We developed a framework that emulates and validates SIMD 2 using NVIDIA GPUs with Tensor Cores. Across 8 applications, SIMD 2 provides up to 38.59× speedup and more than 6.94× on average over optimized CUDA programs, with only 5% of full-chip area overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4932fb48-9654-494d-8e54-1404e09a0a23Cited by top-tier papers2
- Sparsepipe: Sparse Inter-operator Dataflow Architecture with Cross-Iteration ReuseYunan Zhang, Po-An Tsai, Hung-Wei TsengMICRO 2024 · 2 citations
- Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy EfficiencyHansung Kim, Ruohan Richard Yan, Joshua You, Tieliang Vamber Yang et al.ASPLOS 2025 · 1 citation
Builds on14
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingTong Geng, Ang Li, Runbin Shi, Chunshu Wu et al.MICRO 2020 · 299 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi et al.MICRO 2020 · 223 citations
Related papers
- MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsSeunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha et al.MICRO 2023 · 5 citations
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- R2D2: Removing ReDunDancy Utilizing Linearity of Address Generation in GPUsDongho Ha, Yunho Oh, Won Woo RoISCA 2023 · 9 citations
- Carat: Unlocking Value-Level Parallelism for Multiplier-Free GEMMsZhewen Pan, Joshua San Miguel, Di WuASPLOS 2024 · 2 citations
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
