Lune

SC2024Top-tier venue

M3XU: Achieving High-Precision and Complex Matrix Multiplication with Low-Precision MXUs

Dongho Ha, Yunan Zhang, Chen-Chien Kao, Christopher J. Hughes, Won Woo Ro, Hung-Wei Tseng

2024Year
3Citations

Abstract

Beyond the high-profile artificial intelligence and machine learning (AI/ML\mathrm{AI} / \mathrm{ML}) workloads, the demand for high-performance matrix operations on standard and complex floating-point numbers remains strong but underserved. However, the widely adopted low-precision matrix processing units (MXUs) can only fulfill the need for AI/ML workloads, which are underutilized or idle when running applications outside their target domains. This paper presents M3XU\mathbf{M}^{3} \mathbf{X U}, multi-mode matrix processing units that support IEEE 754 single-precision and complex 32bit floating-point numbers. M3XU\mathbf{M}^{3} \mathbf{X U} does not rely on more precise but costly multipliers. Instead, M3XU\mathbf{M}^{3} \mathbf{X U} proposes a multi-step approach that extends existing MXUs for AI/ML workloads. The resulting M3XU\mathbf{M}^{3} \mathbf{X U} can seamlessly upgrade existing systems without programmers’ efforts and maintain the bandwidth demand of existing memory subsystems. This paper evaluates M3XU\mathbf{M}^{3} \mathbf{X U} with full-system emulation and hardware synthesis. M3XU\mathrm{M}^{3} \mathbf{X U} can achieve a 3.64×3.64 \times speedup for 32 -bit matrix multiplications and 3.51×3.51 \times speedup for complex number operations on average compared with conventional vector processing units.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get b3b6991a-e8e0-4439-9057-d0d819bf6890

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines