Efficient Latency-Aware CNN Depth Compression via Two-Stage Dynamic Programming
Jinuk Kim, Yeonwoo Jeong, Deokjae Lee, Hyun Oh Song
摘要
Recent works on neural network pruning advocate that reducing the depth of the network is more effective in reducing run-time memory usage and accelerating inference latency than reducing the width of the network through channel pruning. In this regard, some recent works propose depth compression algorithms that merge convolution layers. However, the existing algorithms have a constricted search space and rely on human-engineered heuristics. In this paper, we propose a novel depth compression algorithm which targets general convolution operations. We propose a subset selection problem that replaces inefficient activation layers with identity functions and optimally merges consecutive convolution operations into shallow equivalent convolution operations for efficient end-to-end inference latency. Since the proposed subset selection problem is NP-hard, we formulate a surrogate optimization problem that can be solved exactly via two-stage dynamic programming within a few seconds. We evaluate our methods and baselines by TensorRT for a fair inference latency comparison. Our method outperforms the baseline method with higher accuracy and faster inference speed in MobileNetV2 on the ImageNet dataset. Specifically, we achieve speed-up with %p accuracy gain in MobileNetV2-1.0 on the ImageNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- Structural Pruning via Latency-Saliency KnapsackMaying Shen, Hongxu Yin, Pavlo Molchanov, Lei Mao 等NeurIPS 2022 · 被引用 70 次
相关 Paper
- Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural NetworksMark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev 等ICML 2020 · 被引用 163 次
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon 等NeurIPS 2024 · 被引用 10 次
- UPSCALE: Unconstrained Channel PruningAlvin Wan, Hanxiang Hao, Kaushik Patnaik, Yueyang Xu 等ICML 2023 · 被引用 7 次
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin 等AAAI 2020 · 被引用 201 次
- Towards Efficient Model Compression via Learned Global RankingTing-Wu Chin, Ruizhou Ding, Cha Zhang, Diana MarculescuCVPR 2020
