Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations
Yichi Zhang, Ritchie Zhao, Weizhe Hua, Nayun Xu, G. Edward Suh, Zhiru Zhang
摘要
We propose precision gating (PG), an end-to-end trainable dynamic dual-precision quantization technique for deep neural networks. PG computes most features in a low precision and only a small proportion of important features in a higher precision to preserve accuracy. The proposed approach is applicable to a variety of DNN architectures and significantly reduces the computational cost of DNN execution with almost no accuracy loss. Our experiments indicate that PG achieves excellent results on CNNs, including statically compressed mobile-friendly networks such as ShuffleNet. Compared to the state-of-the-art prediction-based quantization schemes, PG achieves the same or higher accuracy with 2.4 less compute on ImageNet. PG furthermore applies to RNNs. Compared to 8-bit uniform quantization, PG obtains a 1.2% improvement in perplexity per word with 2.7 computational cost reduction on LSTM on the Penn Tree Bank dataset. Code is available at: this https URL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad 等NeurIPS 2020 · 被引用 149 次
- Dynamic Dual Gating Neural NetworksFanrong Li, Gang Li, Xiangyu He, Jian ChengICCV 2021 · 被引用 40 次
- Prompt-Driven Dynamic Object-Centric Learning for Single Domain GeneralizationDeng Li, Aming Wu, Yaowei Wang, Yahong HanCVPR 2024
它引用的顶会 Paper1
相关 Paper
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang 等DAC 2024 · 被引用 5 次
- AutoQ: Automated Kernel-Wise Neural Network QuantizationQian Lou, Feng Guo, Minje Kim, Lantao Liu 等ICLR 2020 · 被引用 121 次
- Q-PIM: A Genetic Algorithm based Flexible DNN Quantization Method and Application to Processing-In-Memory PlatformYun Long, Edward Lee, Daehyun Kim, Saibal MukhopadhyayDAC 2020 · 被引用 24 次
- DIVISION: Memory Efficient Training via Dual Activation PrecisionGuanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu 等ICML 2023 · 被引用 4 次
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu 等CVPR 2020
