Towards Efficient Model Compression via Learned Global Ranking
Ting-Wu Chin, Ruizhou Ding, Cha Zhang, Diana Marculescu
摘要
Pruning convolutional filters has demonstrated its effectiveness in compressing ConvNets. Prior art in filter pruning requires users to specify a target model complexity (e.g., model size or FLOP count) for the resulting architecture. However, determining a target model complexity can be difficult for optimizing various embodied AI applications such as autonomous robots, drones, and user-facing applications. First, both the accuracy and the speed of ConvNets can affect the performance of the application. Second, the performance of the application can be hard to assess without evaluating ConvNets during inference. As a consequence, finding a sweet-spot between the accuracy and speed via filter pruning, which needs to be done in a trialand-error fashion, can be time-consuming. This work takes a first step toward making this process more efficient by altering the goal of model compression to producing a set of ConvNets with various accuracy and latency trade-offs instead of producing one ConvNet targeting some pre-defined latency constraint. To this end, we propose to learn a global ranking of the filters across different layers of the ConvNet, which is used to obtain a set of ConvNet architectures that have different accuracy/latency trade-offs by pruning the bottom-ranked filters. Our proposed algorithm, LeGR, is shown to be 2× to 3× faster than prior work while having comparable or better performance when targeting seven pruned ResNet-56 with different accuracy/FLOPs profiles on the CIFAR-100 dataset. Additionally, we have evaluated LeGR on ImageNet and Bird-200 with ResNet-50 and Mo-bileNetV2 to demonstrate its effectiveness. Code available at https://github.com/cmu-enyac/LeGR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Structural Pruning for Diffusion ModelsGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2023 · 被引用 257 次
- Pruning Filter in FilterFanxu Meng, Hao Cheng, Ke Li, Huixiang Luo 等NeurIPS 2020 · 被引用 130 次
- CHEX: CHannel EXploration for CNN Model CompressionZejiang Hou, Minghai Qin, Fei Sun, Xiaolong Ma 等CVPR 2022 · 被引用 80 次
- Rethinking the Pruning Criteria for Convolutional Neural NetworkZhongzhan Huang, Wenqi Shao, Xinjiang Wang, Liang Lin 等NeurIPS 2021 · 被引用 75 次
- Structural Pruning via Latency-Saliency KnapsackMaying Shen, Hongxu Yin, Pavlo Molchanov, Lei Mao 等NeurIPS 2022 · 被引用 70 次
它引用的顶会 Paper3
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- Accelerate CNN via Recursive Bayesian PruningYuefu Zhou, Ya Zhang, Yanfeng Wang, Qi TianICCV 2019 · 被引用 64 次
相关 Paper
- Learning Filter Pruning Criteria for Deep Convolutional Neural Networks AccelerationYang He, Yuhang Ding, Ping Liu, Linchao Zhu 等CVPR 2020
- Efficient Latency-Aware CNN Depth Compression via Two-Stage Dynamic ProgrammingJinuk Kim, Yeonwoo Jeong, Deokjae Lee, Hyun Oh SongICML 2023 · 被引用 1 次
- ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksAlexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild 等DAC 2020 · 被引用 5 次
- UPSCALE: Unconstrained Channel PruningAlvin Wan, Hanxiang Hao, Kaushik Patnaik, Yueyang Xu 等ICML 2023 · 被引用 7 次
- Differentiable Neural Network Pruning to Enable Smart Applications on MicrocontrollersEdgar Liberis, Nicholas D. LaneUbiComp 2023 · 被引用 27 次
