Q-PIM: A Genetic Algorithm based Flexible DNN Quantization Method and Application to Processing-In-Memory Platform
Yun Long, Edward Lee, Daehyun Kim, Saibal Mukhopadhyay
Abstract
This paper presents a genetic algorithm (GA) based training free layer-wise quantization method, named as GAQ, to reduce model complexity of arbitrary DNN architectures. The proposed algorithm formulates an optimization problem to determine the quantization level for each DNN layer under the constrain of maximum accuracy degradation and uses genetic algorithm to solve the problem at the inference stage of any pre-trained DNN models. The experimental results on various DNNs for image classification demonstrate 5x to 17x weight compression rate with insignificant (< 2%) accuracy loss, comparable with existing quantization algorithms which typically require multi-pass retraining and handcrafted tuning. To evaluate the computational benefits of GAQ, we present a SRAM based flexible precision all-digital processing-in-memory (PIM) architecture, named as Q-PIM, that leverages GAQ to optimally control precision for each DNN layer to enhance efficiency. The simulation in 28nm CMOS shows potential for significant energy and latency advantage over fixed-precision PIM architectures.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Data-Free Network Compression via Parametric Non-uniform Mixed Precision QuantizationVladimir Chikin, Mikhail AntiukhCVPR 2022 · 17 citations
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao et al.ICLR 2022 · 92 citations
- InfoQ: Mixed-Precision Quantization via Global Information FlowMehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, Manuel RoveriAAAI 2026 · 2 citations
- Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision ActivationsYichi Zhang, Ritchie Zhao, Weizhe Hua, Nayun Xu et al.ICLR 2020 · 28 citations
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
