AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates
Ning Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang, Jian Tang, Jieping Ye
Abstract
Structured weight pruning is a representative model compression technique of DNNs to reduce the storage and computation requirements and accelerate inference. An automatic hyperparameter determination process is necessary due to the large number of flexible hyperparameters. This work proposes AutoCompress, an automatic structured pruning framework with the following key performance improvements: (i) effectively incorporate the combination of structured pruning schemes in the automatic process; (ii) adopt the stateof-art ADMM-based structured weight pruning as the core algorithm, and propose an innovative additional purification step for further weight reduction without accuracy loss; and (iii) develop effective heuristic search method enhanced by experience-based guided search, replacing the prior deep reinforcement learning technique which has underlying incompatibility with the target pruning problem. Extensive experiments on CIFAR-10 and ImageNet datasets demonstrate that AutoCompress is the key to achieve ultra-high pruning rates on the number of weights and FLOPs that cannot be achieved before. As an example, AutoCompress outperforms the prior work on automatic model compression by up to 33× in pruning rate (120× reduction in the actual parameter count) under the same accuracy. Significant inference speedup has been observed from the AutoCompress framework on actual measurements on smartphone. We release all models of this work at anonymous link: http://bit.ly/2VZ63dS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0fa2549-1a3d-41d2-a017-2dbc979b1b0dCited by top-tier papers27
- ResRep: Lossless CNN Pruning via Decoupling Remembering and ForgettingXiaohan Ding, Tianxiang Hao, Jianchao Tan, Ji Liu et al.ICCV 2021 · 202 citations
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-DesignYuxuan Cai, Hongjia Li, Geng Yuan, Wei Niu et al.AAAI 2021 · 124 citations
- Achieving on-Mobile Real-Time Super-Resolution with Neural Architecture and Pruning SearchZheng Zhan, Yifan Gong, Pu Zhao, Geng Yuan et al.ICCV 2021 · 60 citations
- VTC-LFC: Vision Transformer Compression with Low-Frequency ComponentsZhenyu Wang, Hao Luo, Pichao Wang, Feng Ding et al.NeurIPS 2022 · 57 citations
- Co-Modality Graph Contrastive Learning for Imbalanced Node ClassificationYiyue Qian, Chunhui Zhang, Yiming Zhang, Qianlong Wen et al.NeurIPS 2022 · 54 citations
Related papers
- A Unified DNN Weight Pruning Framework Using Reweighted Optimization MethodsTianyun Zhang, Xiaolong Ma, Zheng Zhan, Shanglin Zhou et al.DAC 2021 · 25 citations
- NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile AccelerationZhengang Li, Geng Yuan, Wei Niu, Pu Zhao et al.CVPR 2021
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-Based ApproachHaichuan Yang, Shupeng Gui, Yuhao Zhu, Ji LiuCVPR 2020
- SPDY: Accurate Pruning with Speedup GuaranteesElias Frantar, Dan AlistarhICML 2022 · 45 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
