Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific Computing
Yuechen Lu, Hongwei Zeng, Marc Casas, Weifeng Liu
摘要
Matrix multiplication units (MMUs) in modern parallel processors enable efficient execution of tiled matrix multiplications at varying precisions. While their effectiveness in AI workloads has been well demonstrated, their utility in scientific computing lacks systematic analysis. In this work, we characterize MMUs across a broad range of scientific computing patterns by evaluating performance, power consumption, numerical precision, and memory access behavior. To support this analysis, we develop Cubie, a comprehensive benchmark suite comprising ten MMU-optimized kernels of key parallel patterns. We also categorize MMU utilization patterns into four quadrants and identify the MMU limitations that arise in scientific computing. Through detailed comparisons with vector units, we provide nine key observations on the behavior and implications of MMUs in general scientific workloads, offering valuable insights for architecture, algorithm, and application researchers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
- SparseTIR: Composable Abstractions for Sparse Compilation in Deep LearningZihao Ye, Ruihang Lai, Junru Shao, Tianqi Chen 等ASPLOS 2023 · 被引用 86 次
- TensorIR: An Abstraction for Automatic Tensorized Program OptimizationSiyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin 等ASPLOS 2023 · 被引用 80 次
- QGTC: accelerating quantized graph neural networks via GPU tensor coreYuke Wang, Boyuan Feng, Yufei DingPPoPP 2022 · 被引用 47 次
相关 Paper
- M3XU: Achieving High-Precision and Complex Matrix Multiplication with Low-Precision MXUsDongho Ha, Yunan Zhang, Chen-Chien Kao, Christopher J. Hughes 等SC 2024 · 被引用 3 次
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta 等SC 2024 · 被引用 15 次
- Compiling Strassen-like Matrix Multiplication Algorithms to Fast CUDA KernelsAbhinav JangdaPLDI 2026
- MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsSeunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha 等MICRO 2023 · 被引用 5 次
- A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsMarco Siracusa, Víctor Soria Pardos, Francesco Sgherzi, Joshua Randall 等MICRO 2023 · 被引用 11 次
