Optimizing Information Theory Based Bitwise Bottlenecks for Efficient Mixed-Precision Activation Quantization
Xichuan Zhou, Kui Liu, Cong Shi, Haijun Liu, Ji Liu
摘要
Recent researches on information theory shed new light on the continuous attempts to open the black box of neural signal encoding. Inspired by the problem of lossy signal compression for wireless communication, this paper presents a Bitwise Bottleneck approach for quantizing and encoding neural network activations. Based on the rate-distortion theory, the Bitwise Bottleneck attempts to determine the most significant bits in activation representation by assigning and approximating the sparse coefficients associated with different bits. Given the constraint of a limited average code rate, the bottleneck minimizes the distortion for optimal activation quantization in a flexible layer-by-layer manner. Experiments over ImageNet and other datasets show that, by minimizing the quantization distortion of each layer, the neural network with bottlenecks achieves the state-of-the-art accuracy with low-precision activation. Meanwhile, by reducing the code rate, the proposed method can improve the memory and computational efficiency by over six times compared with the deep neural network with standard single-precision representation. The source code is available on GitHub: https: //github.com/CQUlearningsystemgroup/BitwiseBottleneck .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Wavelet Feature Maps Compression for Image-to-Image CNNsShahaf E. Finder, Yair Zohav, Maor Ashkenazi, Eran TreisterNeurIPS 2022 · 被引用 63 次
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma 等CVPR 2022 · 被引用 38 次
- FQP: A Fibonacci Quantization Processor with Multiplication-Free Computing and Topological-Order RoutingXiaolong Yang, Yang Wang, Yubin Qin, Jiachen Wang 等DAC 2024 · 被引用 1 次
- Information Bottleneck: Exact Analysis of (Quantized) Neural NetworksStephan Sloth Lorenzen, Christian Igel, Mads NielsenICLR 2022 · 被引用 24 次
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
