Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM Inference
Brian Crafton, Xiaochen Peng, Xiaoyu Sun, Ashwin Sanjay Lele, Bo Zhang, Win-San Khwa, Kerem Akarvardar
摘要
To accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve , and TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization FormatsMengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan 等ICML 2026 · 被引用 21 次
- MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model AccelerationSungwoo Kim, Sungbin Kim, Dongho Ha, Hyunwuk Lee 等ISCA 2026
- Is Finer Better? The Limits of Microscaling Formats in Large Language ModelsAndrea Fasoli, Monodeep Kar, Chi-Chun (Charlie) Liu, Swagath Venkataramani 等ICLR 2026 · 被引用 7 次
- An Algorithm-Hardware Co-design Based on Revised Microscaling Format Quantization for Accelerating Large Language ModelsYingbo Hao, Huangxu Chen, Yi Zou, Yanfeng YangDAC 2025 · 被引用 1 次
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error ReductionJatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan 等ICML 2026 · 被引用 6 次
