SC2024Top-tier venue
M3XU: Achieving High-Precision and Complex Matrix Multiplication with Low-Precision MXUs
Dongho Ha, Yunan Zhang, Chen-Chien Kao, Christopher J. Hughes, Won Woo Ro, Hung-Wei Tseng
Abstract
Beyond the high-profile artificial intelligence and machine learning () workloads, the demand for high-performance matrix operations on standard and complex floating-point numbers remains strong but underserved. However, the widely adopted low-precision matrix processing units (MXUs) can only fulfill the need for AI/ML workloads, which are underutilized or idle when running applications outside their target domains. This paper presents , multi-mode matrix processing units that support IEEE 754 single-precision and complex 32bit floating-point numbers. does not rely on more precise but costly multipliers. Instead, proposes a multi-step approach that extends existing MXUs for AI/ML workloads. The resulting can seamlessly upgrade existing systems without programmers’ efforts and maintain the bandwidth demand of existing memory subsystems. This paper evaluates with full-system emulation and hardware synthesis. can achieve a speedup for 32 -bit matrix multiplications and speedup for complex number operations on average compared with conventional vector processing units.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b3b6991a-e8e0-4439-9057-d0d819bf6890Related papers
- Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific ComputingYuechen Lu, Hongwei Zeng, Marc Casas, Weifeng LiuPPoPP 2026 · 1 citation
- MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsSeunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha et al.MICRO 2023 · 5 citations
- MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model AccelerationSungwoo Kim, Sungbin Kim, Dongho Ha, Hyunwuk Lee et al.ISCA 2026
- SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMMYunan Zhang, Po-An Tsai, Hung-Wei TsengISCA 2022 · 6 citations
- HiT: A Unified Sparsity-Adaptive Architecture for High-Throughput Matrix MultiplicationTingting Xiang, Xiaochen Wang, Miao Yu, Trevor E. CarlsonISCA 2026
