Lune

SC2024顶会

M3XU: Achieving High-Precision and Complex Matrix Multiplication with Low-Precision MXUs

Dongho Ha, Yunan Zhang, Chen-Chien Kao, Christopher J. Hughes, Won Woo Ro, Hung-Wei Tseng

2024年份
3被引次数

摘要

Beyond the high-profile artificial intelligence and machine learning (AI/ML\mathrm{AI} / \mathrm{ML}) workloads, the demand for high-performance matrix operations on standard and complex floating-point numbers remains strong but underserved. However, the widely adopted low-precision matrix processing units (MXUs) can only fulfill the need for AI/ML workloads, which are underutilized or idle when running applications outside their target domains. This paper presents M3XU\mathbf{M}^{3} \mathbf{X U}, multi-mode matrix processing units that support IEEE 754 single-precision and complex 32bit floating-point numbers. M3XU\mathbf{M}^{3} \mathbf{X U} does not rely on more precise but costly multipliers. Instead, M3XU\mathbf{M}^{3} \mathbf{X U} proposes a multi-step approach that extends existing MXUs for AI/ML workloads. The resulting M3XU\mathbf{M}^{3} \mathbf{X U} can seamlessly upgrade existing systems without programmers’ efforts and maintain the bandwidth demand of existing memory subsystems. This paper evaluates M3XU\mathbf{M}^{3} \mathbf{X U} with full-system emulation and hardware synthesis. M3XU\mathrm{M}^{3} \mathbf{X U} can achieve a 3.64×3.64 \times speedup for 32 -bit matrix multiplications and 3.51×3.51 \times speedup for complex number operations on average compared with conventional vector processing units.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖