Lune

DAC2025顶会

Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM Inference

Brian Crafton, Xiaochen Peng, Xiaoyu Sun, Ashwin Sanjay Lele, Bo Zhang, Win-San Khwa, Kerem Akarvardar

2025年份
2被引次数

摘要

To accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of 35,000\mathbf{3 5, 0 0 0} MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve 1.8×2.2×1.8 \times 2.2 \times, and 1.9×1.9 \times TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖