Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
Sung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi, Hayden K. H. So, Xuehai Qian, Yanzhi Wang, Xue Lin
Abstract
Deep Neural Networks (DNNs) have achieved extraordinary performance in various application domains. To support diverse DNN models, efficient implementations of DNN inference on edge-computing platforms, e.g., ASICs, FPGAs, and embedded systems, are extensively investigated. Due to the huge model size and computation amount, model compression is a critical step to deploy DNN models on edge devices. This paper focuses on weight quantization, a hardware-friendly model compression approach that is complementary to weight pruning.
Unlike existing methods that use the same quantization scheme for all weights, we propose the first solution that applies different quantization schemes for different rows of the weight matrix. It is motivated by (1) the distribution of the weights in the different rows are not the same; and (2) the potential of achieving better utilization of heterogeneous FPGA hardware resources. To achieve that, we first propose a hardware-friendly quantization scheme named sum-of-power-of-2 (SP2) suitable for Gaussianlike weight distribution, in which the multiplication arithmetic can be replaced with logic shifter and adder, thereby enabling highly efficient implementations with the FPGA LUT resources. In contrast, the existing fixed-point quantization is suitable for Uniform-like weight distribution and can be implemented efficiently by DSP. Then to fully explore the resources, we propose an FPGA-centric mixed scheme quantization (MSQ) with an ensemble of the proposed SP2 and the fixed-point schemes.
Combining the two schemes can maintain, or even increase accuracy due to better matching with weight distributions.
For the FPGA implementations, we develop a parameterized architecture with heterogeneous Generalized Matrix Multiplication (GEMM) cores-one using LUTs for computations with SP2 quantized weights and the other utilizing DSPs for fixedpoint quantized weights. Given the partition ratio among the two schemes based on resource characterization, MSQ quantization training algorithm derives an optimally quantized model for the FPGA implementation. We evaluate our FPGA-centric quantization framework across multiple application domains. With optimal SP2/fixed-point ratios on two FPGA devices, i.e., Zynq XC7Z020 and XC7Z045, we achieve performance improvement of 2.1 × -4.1× compared to solely exploiting DSPs for all multiplication operations. In addition, the CNN implementations with the proposed MSQ scheme can achieve higher accuracy and comparable hardware utilization efficiency compared to the state-of-the-art designs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 480b8fda-448f-43ff-afa1-909b33f9daa1Cited by top-tier papers5
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationMinsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi et al.ISCA 2025 · 17 citations
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple PrecisionsSung-En Chang, Yanyu Li, Mengshu Sun, Weiwen Jiang et al.ICCV 2021 · 14 citations
- EBSP: evolving bit sparsity patterns for hardware-friendly inference of quantized deep neural networksFangxin Liu, Wenbo Zhao, Zongwu Wang, Yongbiao Chen et al.DAC 2022 · 12 citations
- Be Like Water: Adaptive Floating Point for Machine LearningThomas Y. Yeh, Max Sterner, Zerlina Lai, Brandon Chuang et al.ICML 2022 · 11 citations
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park et al.HPCA 2025 · 9 citations
Builds on4
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
- Towards Unified INT8 Training for Convolutional Neural NetworkFeng Zhu, Ruihao Gong, Fengwei Yu, Xianglong Liu et al.CVPR 2020
Related papers
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 28 citations
- MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector QuantizationShuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang et al.ASPLOS 2025 · 5 citations
- BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNsYongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim et al.SC 2020 · 31 citations
- MSQ: Memory-Efficient Bit Sparsification QuantizationSeokho Han, Seoyeon Yoon, Jinhee Kim, Dongwei Wang et al.ICCV 2025 · 2 citations
- FATE: Boosting the Performance of Hyper-Dimensional Computing Intelligence with Flexible Numerical DAta TypEHaomin Li, Fangxin Liu, Yichi Chen, Zongwu Wang et al.ISCA 2025 · 4 citations
