PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile Devices
Xiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin, Jian Tang, Kaisheng Ma, Bin Ren, Yanzhi Wang
摘要
Model compression techniques on Deep Neural Network (DNN) have been widely acknowledged as an effective way to achieve acceleration on a variety of platforms, and DNN weight pruning is a straightforward and effective method. There are currently two mainstreams of pruning methods representing two extremes of pruning regularity: non-structured, fine-grained pruning can achieve high sparsity and accuracy, but is not hardware friendly; structured, coarse-grained pruning exploits hardware-efficient structures in pruning, but suffers from accuracy drop when the pruning rate is high. In this paper, we introduce PCONV, comprising a new sparsity dimension, – fine-grained pruning patterns inside the coarse-grained structures. PCONV comprises two types of sparsities, Sparse Convolution Patterns (SCP) which is generated from intra-convolution kernel pruning and connectivity sparsity generated from inter-convolution kernel pruning. Essentially, SCP enhances accuracy due to its special vision properties, and connectivity sparsity increases pruning rate while maintaining balanced workload on filter computation. To deploy PCONV, we develop a novel compiler-assisted DNN inference framework and execute PCONV models in real-time without accuracy compromise, which cannot be achieved in prior work. Our experimental results show that, PCONV outperforms three state-of-art end-to-end DNN frameworks, TensorFlow-Lite, TVM, and Alibaba Mobile Neural Network with speedup up to 39.2 ×, 11.4 ×, and 6.3 ×, respectively, with no accuracy loss. Mobile devices can achieve real-time inference on large-scale DNNs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang 等ASPLOS 2020 · 被引用 214 次
- YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-DesignYuxuan Cai, Hongjia Li, Geng Yuan, Wei Niu 等AAAI 2021 · 被引用 124 次
- MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the EdgeGeng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li 等NeurIPS 2021 · 被引用 124 次
- Advancing Model Pruning via Bi-level OptimizationYihua Zhang, Yuguang Yao, Parikshit Ram, Pu Zhao 等NeurIPS 2022 · 被引用 101 次
- SparCL: Sparse Continual Learning on the EdgeZifeng Wang, Zheng Zhan, Yifan Gong, Geng Yuan 等NeurIPS 2022 · 被引用 97 次
相关 Paper
- RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesWei Niu, Mengshu Sun, Zhengang Li, Jou-An Chen 等AAAI 2021 · 被引用 14 次
- DEPrune: Depth-wise Separable Convolution Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyunchan Moon, Myung Kuk Yoon 等NeurIPS 2024 · 被引用 10 次
- PCNN: Pattern-based Fine-Grained Regular Pruning Towards Optimizing CNN AcceleratorsZhanhong Tan, Jiebo Song, Xiaolong Ma, Sia Huat Tan 等DAC 2020 · 被引用 28 次
- PENNI: Pruned Kernel Sharing for Efficient CNN InferenceShiyu Li, Edward Hanson, Hai Li, Yiran ChenICML 2020 · 被引用 23 次
- High Performance Depthwise and Pointwise Convolutions on Mobile DevicesPengfei Zhang, Eric Lo, Baotong LuAAAI 2020 · 被引用 54 次
