SC2025Top-tier venue
MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library
Weihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng
2025Year
2Citations
Abstract
Micro-scaling General Matrix Multiplication (MX-GEMM), which leverages 8-bit micro-scaling format (MX-format) inputs, represents a significant step forward in accelerating deep learning workloads. The MX-format space is diverse, encompassing various scaling patterns and granularities. However, current MX-GEMM implementations typically adopt a model-oriented approach, where format customization is tailored to individual models. This results in three key limitations: rigid problem-kernel coupling, inefficient promotion operations, and overlooked quantization overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- An Algorithm-Hardware Co-design Based on Revised Microscaling Format Quantization for Accelerating Large Language ModelsYingbo Hao, Huangxu Chen, Yi Zou, Yanfeng YangDAC 2025 · 1 citation
- Avant-Garde: Empowering GPUs with Scaled Numeric FormatsMinseong Gil, Dongho Ha, Simla Burcu Harma, Myung Kuk Yoon et al.ISCA 2025 · 4 citations
- Is Finer Better? The Limits of Microscaling Formats in Large Language ModelsAndrea Fasoli, Monodeep Kar, Chi-Chun (Charlie) Liu, Swagath Venkataramani et al.ICLR 2026 · 7 citations
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error ReductionJatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan et al.ICML 2026 · 6 citations
- MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsSeunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha et al.MICRO 2023 · 5 citations
