Avant-Garde: Empowering GPUs with Scaled Numeric Formats
Minseong Gil, Dongho Ha, Simla Burcu Harma, Myung Kuk Yoon, Babak Falsafi, Won Woo Ro, Yunho Oh
摘要
The escalating computational and memory demands of deep neural networks have outpaced chip density improvements, making arithmetic density a key bottleneck for GPUs.Scaled numeric formats, such as FP8 and Microscaling (MX), improve arithmetic density by applying adaptive scaling factors across varying block sizes and multiple scaling hierarchies.Unfortunately, supporting diverse scaled numeric formats often requires GPUs to rely on softwarebased implementations, increasing instruction and register overhead and degrading performance.We propose Avant-Garde, a GPU microarchitecture that natively supports diverse scaled numeric formats by converting them into a consistent single-level internal representation.Avant-Garde integrates an Operand Transformer, a hardware module that dynamically flattens multi-level scaling formats into single-level internal representations, a novel Tensor Core, and an optimized data layout to eliminate instruction and register overhead.Our evaluations show that Avant-Garde achieves up to 74% higher throughput and 44% lower execution time, while maintaining accuracy within 0.2% compared to conventional GPUs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationWeiming Hu, Zihan Zhang, Haoyan Zhang, Chen Zhang 等ASPLOS 2026 · 被引用 2 次
- Hardwired-Neuron Language Processing Units as General-Purpose Cognitive SubstratesYang Liu, Yi Chen, Yongwei Zhao, Yifan Hao 等ASPLOS 2026
- BCCE: Block-Centric GPU Co-Design for Real-Time Range-Top-K Query at ScaleChengying Huan, Ziheng Meng, Zhengyi Yang, Yongchao Liu 等HPDC 2026
相关 Paper
- MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model AccelerationSungwoo Kim, Sungbin Kim, Dongho Ha, Hyunwuk Lee 等ISCA 2026
- Is Finer Better? The Limits of Microscaling Formats in Large Language ModelsAndrea Fasoli, Monodeep Kar, Chi-Chun (Charlie) Liu, Swagath Venkataramani 等ICLR 2026 · 被引用 7 次
- MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM LibraryWeihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 等SC 2025 · 被引用 2 次
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error ReductionJatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan 等ICML 2026 · 被引用 6 次
- Unit Scaling: Out-of-the-Box Low-Precision TrainingCharlie Blake, Douglas Orr, Carlo LuschiICML 2023 · 被引用 15 次
