Avant-Garde: Empowering GPUs with Scaled Numeric Formats
Minseong Gil, Dongho Ha, Simla Burcu Harma, Myung Kuk Yoon, Babak Falsafi, Won Woo Ro, Yunho Oh
Abstract
The escalating computational and memory demands of deep neural networks have outpaced chip density improvements, making arithmetic density a key bottleneck for GPUs.Scaled numeric formats, such as FP8 and Microscaling (MX), improve arithmetic density by applying adaptive scaling factors across varying block sizes and multiple scaling hierarchies.Unfortunately, supporting diverse scaled numeric formats often requires GPUs to rely on softwarebased implementations, increasing instruction and register overhead and degrading performance.We propose Avant-Garde, a GPU microarchitecture that natively supports diverse scaled numeric formats by converting them into a consistent single-level internal representation.Avant-Garde integrates an Operand Transformer, a hardware module that dynamically flattens multi-level scaling formats into single-level internal representations, a novel Tensor Core, and an optimized data layout to eliminate instruction and register overhead.Our evaluations show that Avant-Garde achieves up to 74% higher throughput and 44% lower execution time, while maintaining accuracy within 0.2% compared to conventional GPUs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang et al.ASPLOS 2026 · 2 citations
- Hardwired-Neuron Language Processing Units as General-Purpose Cognitive SubstratesYang Liu, Yi Chen, Yongwei Zhao, Yifan Hao et al.ASPLOS 2026
- BCCE: Block-Centric GPU Co-Design for Real-Time Range-Top-K Query at ScaleChengying Huan, Ziheng Meng, Zhengyi Yang, Yongchao Liu et al.HPDC 2026
Related papers
- MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model AccelerationSungwoo Kim, Sungbin Kim, Dongho Ha, Hyunwuk Lee et al.ISCA 2026
- Is Finer Better? The Limits of Microscaling Formats in Large Language ModelsAndrea Fasoli, Monodeep Kar, Chi-Chun (Charlie) Liu, Swagath Venkataramani et al.ICLR 2026 · 7 citations
- MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM LibraryWeihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou et al.SC 2025 · 2 citations
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error ReductionJatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan et al.ICML 2026 · 6 citations
- Unit Scaling: Out-of-the-Box Low-Precision TrainingCharlie Blake, Douglas Orr, Carlo LuschiICML 2023 · 15 citations
