Data-Free Network Compression via Parametric Non-uniform Mixed Precision Quantization
Vladimir Chikin, Mikhail Antiukh
摘要
Deep Neural Networks (DNNs) usually have a large number of parameters and consume a huge volume of storage space, which limits the application of DNNs on memory-constrained devices. Network quantization is an appealing way to compress DNNs. However, most of existing quantization methods require the training dataset and a fine-tuning procedure to preserve the quality of a full-precision model. These are unavailable for the confidential scenarios due to personal privacy and security problems. Focusing on this issue, we propose a novel data-free method for network compression called PNMQ, which employs the Parametric Non-uniform Mixed precision Quantization to generate a quantized network. During the compression stage, the optimal parametric non-uniform quantization grid is calculated directly for each layer to minimize the quantization error. User can directly specify the required compression ratio of a network, which is used by the PNMQ algorithm to select bitwidths of layers. This method does not require any model retraining or expensive calculations, which allows efficient implementations for network compression on edge devices. Extensive experiments have been conducted on various computer vision tasks and the results demonstrate that PNMQ achieves better performance than other state-of-the-art methods of network compression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Lossy and Lossless (L2) Post-training Model Size CompressionYumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong 等ICCV 2023 · 被引用 5 次
- AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything ModelWenlun Zhang, Yunshan Zhong, Shimpei Ando, Kentaro YoshiokaICCV 2025 · 被引用 3 次
它引用的顶会 Paper4
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- The Knowledge Within: Methods for Data-Free Model CompressionMatan Haroush, Itay Hubara, Elad Hoffer, Daniel SoudryCVPR 2020
相关 Paper
- Unified Data-Free Compression: Pruning and Quantization without Fine-TuningShipeng Bai, Jun Chen, Xintian Shen, Yixuan Qian 等ICCV 2023 · 被引用 31 次
- PowerQuant: Automorphism Search for Non-Uniform QuantizationEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyICLR 2023 · 被引用 1 次
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao 等ICLR 2022 · 被引用 92 次
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-Based ApproachHaichuan Yang, Shupeng Gui, Yuhao Zhu, Ji LiuCVPR 2020
- Causal-DFQ: Causality Guided Data-free Network QuantizationYuzhang Shang, Bingxin Xu, Gaowen Liu, Ramana Rao Kompella 等ICCV 2023 · 被引用 8 次
