A2Q: Accumulator-Aware Quantization with Guaranteed Overflow Avoidance
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig
摘要
We present accumulator-aware quantization (A2Q), a novel weight quantization method designed to train quantized neural networks (QNNs) to avoid overflow when using low-precision accumulators during inference. A2Q introduces a unique formulation inspired by weight normalization that constrains the ℓ 1 -norm of model weights according to accumulator bit width bounds that we derive. Thus, in training QNNs for low-precision accumulation, A2Q also inherently promotes unstructured weight sparsity to guarantee overflow avoidance. We apply our method to deep learning-based computer vision tasks to show that A2Q can train QNNs for low-precision accumulators while maintaining model accuracy competitive with a floating-point baseline. In our evaluations, we consider the impact of A2Q on both general-purpose platforms and programmable hardware. However, we primarily target model deployment on FPGAs because they can be programmed to fully exploit custom accumulator bit widths. Our experimentation shows accumulator bit width significantly impacts the resource efficiency of FPGA-based accelerators. On average across our benchmarks, A2Q offers up to a 2.3x reduction in resource utilization over 32-bit accumulator counterparts with 99.2% of the floating-point model accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A2Q+: Improving Accumulator-Aware Weight QuantizationIan Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig, Yaman UmurogluICML 2024 · 被引用 11 次
- PsumQuant: In-line Post-training Partial Sum Quantizer for Energy Efficient NPU InferenceSangwoo Hwang, Yeeun Hong, Jaeha KungICML 2026
它引用的顶会 Paper7
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 被引用 315 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
相关 Paper
- Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsYaniv Blumenfeld, Itay Hubara, Daniel SoudryICLR 2024 · 被引用 5 次
- WrapNet: Neural Net Inference with Ultra-Low-Precision ArithmeticRenkun Ni, Hong-Min Chu, Oscar Castañeda, Ping-yeh Chiang 等ICLR 2021 · 被引用 16 次
- Learnable Companding Quantization for Accurate Low-Bit Neural NetworksKohei YamamotoCVPR 2021
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksYikai Wang, Yi Yang, Fuchun Sun, Anbang YaoICCV 2021 · 被引用 18 次
- Cambricon-Q: A Hybrid Architecture for Efficient TrainingYongwei Zhao, Chang Liu, Zidong Du, Qi Guo 等ISCA 2021 · 被引用 28 次
