Lune

DAC2025顶会

An Algorithm-Hardware Co-design Based on Revised Microscaling Format Quantization for Accelerating Large Language Models

Yingbo Hao, Huangxu Chen, Yi Zou, Yanfeng Yang

2025年份
1被引次数

摘要

The narrow-bit-width data format is crucial for reducing the computation and storage costs of modern deep learning applications, particularly in large language models (LLMs) based applications. Microscaling (MX) format has been proven as a drop-in replacement for the baseline FP32 in existing inference frameworks, with low user friction. However, deploying such a new format into existing hardware systems is still challenging, and the dominant solution for LLM inference at low precision is still low-bit quantization. This particularly limits the strategic applications of such LLMs in real deployment on a large scale. In this work, we propose an algorithm-hardware co-design that adopts a two-level RevisedMXFormatQuantizationRevised MX Format Quantization (RMFQ) and a RevisedMXFormatAcceleratorRevised MX Format Accelerator (RMFA) architecture design. RMFQ proposes the revisedMX(RMX)revised M X(R M X) format and provides a novel quantization framework with innovative group direction. Also, RMFA provides an RMX adaptive hardware architecture and an RMX encoding scheme. As a result, RMFQ pushes the limit of 4-bit and 6-bit quantization to a new state-of-the-art, and RMFA surpasses the existing outlier-aware accelerator such as OliVe, achieving a 1.28×1.28 \times speedup and a 1.31×1.31 \times energy reduction.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖