NPAS: A Compiler-Aware Framework of Unified Network Pruning and Architecture Search for Beyond Real-Time Mobile Acceleration
Zhengang Li, Geng Yuan, Wei Niu, Pu Zhao, Yanyu Li, Yuxuan Cai, Xuan Shen, Zheng Zhan, Zhenglun Kong, Qing Jin, Zhiyu Chen, Sijia Liu
摘要
With the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Prior methods towards this goal, including model compression and network architecture search (NAS), are largely performed independently, and do not fully consider compiler-level optimizations which is a must-do for mobile acceleration. In this work, we first propose (i) a general category of fine-grained structured pruning applicable to various DNN layers, and (ii) a comprehensive, compiler automatic code generation framework supporting different DNNs and different pruning schemes, which bridge the gap of model compression and NAS. We further propose NPAS, a compiler-aware unified network pruning and architecture search. To deal with large search space, we propose a meta-modeling procedure based on reinforcement learning with fast evaluation and Bayesian optimization, ensuring the total number of training epochs comparable with representative NAS frameworks. Our framework achieves 6.7ms, 5.9ms, and 3.9ms ImageNet inference times with 78.2%, 75% (MobileNet-V3 level), and 71% (MobileNet-V2 level) Top-1 accuracy respectively on an off-the-shelf mobile phone, consistently outperforming prior work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante 等ICLR 2022 · 被引用 57 次
- Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement LearningSixing Yu, Arya Mazaheri, Ali JannesariICML 2022 · 被引用 54 次
- Balanced Column-Wise Block Pruning for Maximizing GPU ParallelismCheonjun Park, Mincheol Park, Hyun Jae Oh, Minkyu Kim 等AAAI 2023 · 被引用 14 次
- Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile DevicesZhihao Shu, Xiaowei Yu, Zihao Wu, Wenqi Jia 等NeurIPS 2024 · 被引用 6 次
- Interspace Pruning: Using Adaptive Filter Representations to Improve Training of Sparse CNNsPaul Wimmer, Jens Mehnert, Alexandru ConduracheCVPR 2022
它引用的顶会 Paper7
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- FairNAS: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture SearchXiangxiang Chu, Bo Zhang, Ruijun XuICCV 2021 · 被引用 362 次
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang 等ASPLOS 2020 · 被引用 214 次
- AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression RatesNing Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang 等AAAI 2020 · 被引用 204 次
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin 等AAAI 2020 · 被引用 201 次
相关 Paper
- Neural Pruning Search for Real-Time Object Detection of Autonomous VehiclesPu Zhao, Geng Yuan, Yuxuan Cai, Wei Niu 等DAC 2021 · 被引用 23 次
- Auto Graph Encoder-Decoder for Neural Network PruningSixing Yu, Arya Mazaheri, Ali JannesariICCV 2021 · 被引用 47 次
- AutoShrink: A Topology-Aware NAS for Discovering Efficient Neural ArchitectureTunhou Zhang, Hsin-Pai Cheng, Zhenwen Li, Feng Yan 等AAAI 2020 · 被引用 9 次
- RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesWei Niu, Mengshu Sun, Zhengang Li, Jou-An Chen 等AAAI 2021 · 被引用 14 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
