AMPA: Adaptive Mixed Precision Allocation for Low-Bit Integer Training
Li Ding, Wen Fei, Yuyang Huang, Shuangrui Ding, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong
摘要
Low-bit integer training emerges as a promising approach to mitigate the heavy burden during network training by quantizing the weights, activations, and gradients. However, existing methods cannot well achieve mixed-precision quantization for low-bit training and are commonly limited to INT8 precision. In this paper, we propose a novel low-bit integer training framework that, for the first time, achieves adaptive mixed-precision allocation (AMPA) for weights, activations, and gradients, and pushes the boundaries to a precision level below INT8. We develop a novel magnitude-based sensitivity measurement with regard to the quantization losses of weight, activation, and gradient quantization and the average gradient magnitudes, which is demonstrated as an upper bound of quantization influence in theory. We further design a layer-wise precision update strategy under observations on the quantization losses and their effects on model performance in low-bit training. Extensive experiments on different backbones and datasets show that, compared to INT8 quantization, the proposed method can achieve more than 38% BitOPs reduction with a tolerable loss below 2% in image classification, image segmentation, and language modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami 等NeurIPS 2020 · 被引用 434 次
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami 等ICML 2021 · 被引用 240 次
- Learned Token Pruning for TransformersSehoon Kim, Sheng Shen, David Thorsley, Amir Gholami 等KDD 2022 · 被引用 97 次
相关 Paper
- ALAM: Averaged Low-Precision Activation for Memory-Efficient Training of Transformer ModelsSunghyeon Woo, Sunwoo Lee, Dongsuk JeonICLR 2024 · 被引用 4 次
- InfoQ: Mixed-Precision Quantization via Global Information FlowMehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, Manuel RoveriAAAI 2026 · 被引用 2 次
- Distribution Adaptive INT8 Quantization for Training CNNsKang Zhao, Sida Huang, Pan Pan, Yinghan Li 等AAAI 2021 · 被引用 86 次
- Double Rounding: Nearly Lossless Adaptive Bit Switching for QATHaiduo Huang, Zhenhua Liu, Tian Xia, Pengju RenAAAI 2026
- HLHLp: Quantized Neural Networks Training for Reaching Flat Minima in Loss SurfaceSungho Shin, Jinhwan Park, Yoonho Boo, Wonyong SungAAAI 2020 · 被引用 6 次
