PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural Networks
Marina Neseem, Conor McCullough, Randy Hsin, Chas Leichner, Shan Li, In Suk Chong, Andrew G. Howard, Lukasz Lew, Sherief Reda, Ville-Mikko Rautio, Daniele Moro
Abstract
Low-precision quantization is recognized for its efficacy in neural network optimization. Our analysis reveals that non-quantized elementwise operations which are prevalent in layers such as parameterized activation functions, batch normalization, and quantization scaling dominate the inference cost of low-precision models. These non-quantized elementwise operations are commonly overlooked in SOTA efficiency metrics such as Arithmetic Computation Effort (ACE) [46]. In this paper, we propose ACE<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">v2</inf> - an extended version of ACE which offers a better alignment with the inference cost of quantized models and their energy consumption on ML hardware. Moreover, we introduce PikeLPN<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Pike is a slim fast fish, LPN stands for Low-Precision Network., a model that addresses these efficiency issues by applying quantization to both elementwise operations and multiply-accumulate operations. In particular, we present a novel quantization technique for batch normalization layers named QuantNorm which allows for quantizing the batch normalization parameters without compromising the model performance. Additionally, we propose applying Double Quantization where the quantization scaling parameters are quantized. Furthermore, we recognize and resolve the issue of distribution mismatch in Separable Convolution layers by introducing Distribution-Heterogeneous Quantization which enables quantizing them to low-precision. PikeLPN achieves Pareto-optimality in efficiency-accuracy trade-off with up to 3× efficiency improvement compared to SOTA low-precision models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1733139b-9022-476e-bdc5-e5d3a8ec5859Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 315 citations
- How Do Adam and Training Strategies Help BNNs OptimizationZechun Liu, Zhiqiang Shen, Shichao Li, Koen Helwegen et al.ICML 2021 · 100 citations
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- PokeBNN: A Binary Pursuit of Lightweight AccuracyYichi Zhang, Zhiru Zhang, Lukasz LewCVPR 2022 · 38 citations
- Binarizing MobileNet via Evolution-Based SearchingHai Phan, Zechun Liu, Dang Huynh, Marios Savvides et al.CVPR 2020
Related papers
- Linear Symmetric Quantization of Neural Networks for Low-precision Integer HardwareXiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu et al.ICLR 2020 · 67 citations
- Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice SparsityDongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong et al.HPCA 2025 · 5 citations
- Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN InferenceAkshat Ramachandran, Zishen Wan, Geonhwa Jeong, John L. Gustafson et al.DAC 2024 · 17 citations
- FQP: A Fibonacci Quantization Processor with Multiplication-Free Computing and Topological-Order RoutingXiaolong Yang, Yang Wang, Yubin Qin, Jiachen Wang et al.DAC 2024 · 1 citation
- Towards Mixed-Precision Quantization of Neural Networks via Constrained OptimizationWeihan Chen, Peisong Wang, Jian ChengICCV 2021 · 91 citations
