Linear Symmetric Quantization of Neural Networks for Low-precision Integer Hardware
Xiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu, Lei Zhang
Abstract
With the proliferation of specialized neural network processors that operate on low-precision integers, the performance of Deep Neural Network inference becomes increasingly dependent on the result of quantization. Despite plenty of prior work on the quantization of weights or activations for neural networks, there is still a wide gap between the software quantizers and the low-precision accelerator implementation, which degrades either the efficiency of networks or that of the hardware for the lack of software and hardware coordination at design-phase. In this paper, we propose a learned linear symmetric quantizer for integer neural network processors, which not only quantizes neural parameters and activations to low-bit integer but also accelerates hardware inference by using batch normalization fusion and low-precision accumulators (e.g., 16-bit) and multipliers (e.g., 4-bit). We use a unified way to quantize weights and activations, and the results outperform many previous approaches for various networks such as AlexNet, ResNet, and lightweight models like MobileNet while keeping friendly to the accelerator architecture. Additional, we also apply the method to object detection models and witness high performance and accuracy in YOLO-v2. Finally, we deploy the quantized models on our specialized integer-arithmetic-only DNN accelerator to show the effectiveness of the proposed quantizer. We show that even with linear symmetric quantization, the results can be better than asymmetric or non-linear methods in 4-bit networks. In evaluation, the proposed quantizer induces less than 0.4% accuracy drop in ResNet18, ResNet34, and AlexNet when quantizing the whole network as required by the integer processors.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 14d638ef-7ebd-410c-8ab0-8265d3e7330fCited by top-tier papers15
- Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer QuantizationJeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park et al.NeurIPS 2023 · 157 citations
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through EstimationZechun Liu, Kwang-Ting Cheng, Dong Huang, Eric P. Xing et al.CVPR 2022 · 108 citations
- FlexRound: Learnable Rounding based on Element-wise Division for Post-Training QuantizationJung Hyun Lee, Jeonghoon Kim, Se Jung Kwon, Dongsoo LeeICML 2023 · 56 citations
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 45 citations
- Improving Low-Precision Network Quantization via Bin RegularizationTiantian Han, Dong Li, Ji Liu, Lu Tian et al.ICCV 2021 · 45 citations
Related papers
- Learnable Companding Quantization for Accurate Low-Bit Neural NetworksKohei YamamotoCVPR 2021
- BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionFangxin Liu, Ning Yang, Zongwu Wang, Xuanpeng Zhu et al.DAC 2025 · 1 citation
- Precon: A Precision-Convertible Architecture for Accelerating Quantized Deep Learning Models across Various Domains Including LLMsJongwoo Park, Hyeonseong Kim, Jiyun Han, Seungkyu ChoiDAC 2025
- WrapNet: Neural Net Inference with Ultra-Low-Precision ArithmeticRenkun Ni, Hong-Min Chu, Oscar Castañeda, Ping-yeh Chiang et al.ICLR 2021 · 16 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
