APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
Tianzhe Wang, Kuan Wang, Han Cai, Ji Lin, Zhijian Liu, Hanrui Wang, Yujun Lin, Song Han
摘要
We present APQ, a novel design methodology for efficient deep learning deployment. Unlike previous methods that separately optimize the neural network architecture, pruning policy, and quantization policy, we design to optimize them in a joint manner. To deal with the larger design space it brings, we devise to train a quantizationaware accuracy predictor that is fed to the evolutionary search to select the best fit. Since directly training such a predictor requires time-consuming quantization data collection, we propose to use predictor-transfer technique to get the quantization-aware predictor: we first generate a large dataset of NN architecture, ImageNet accuracy pairs by sampling a pretrained unified once-for-all network and doing direct evaluation; then we use these data to train an accuracy predictor without quantization, followed by transferring its weights to train the quantization-aware predictor, which largely reduces the quantization data collection time. Extensive experiments on ImageNet show the benefits of this joint design methodology: the model searched by our method maintains the same level accuracy as ResNet34 8-bit model while saving 8× BitOps; we achieve 2×/1.3× latency/energy saving compared to 36] while obtaining the same level accuracy; the marginal search cost of joint optimization for a new deployment scenario outperforms separate optimizations using ProxylessNAS+AMC+HAQ [5, 12, 36] by 2.3% accuracy while reducing orders of magnitude GPU hours and CO 2 emission with respect to the training cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 被引用 412 次
- EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense PredictionHan Cai, Junyan Li, Muyan Hu, Chuang Gan 等ICCV 2023 · 被引用 265 次
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
- QuantumNAS: Noise-Adaptive Search for Robust Quantum CircuitsHanrui Wang, Yongshan Ding, Jiaqi Gu, Yujun Lin 等HPCA 2022 · 被引用 199 次
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly 等AAAI 2021 · 被引用 79 次
它引用的顶会 Paper3
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
相关 Paper
- SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 InferenceXudong Wang, Li Lyna Zhang, Jiahang Xu, Quanlu Zhang 等ICCV 2023 · 被引用 3 次
- Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture SearchMingzhu Shen, Feng Liang, Ruihao Gong, Yuhang Li 等ICCV 2021 · 被引用 50 次
- JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-ExplorationMingzi Wang, Yuan Meng, Chen Tang, Weixiang Zhang 等AAAI 2025 · 被引用 3 次
- Hardware-adaptive Efficient Latency Prediction for NAS via Meta-LearningHayeon Lee, Sewoong Lee, Song Chong, Sung Ju HwangNeurIPS 2021 · 被引用 32 次
- CompOFA - Compound Once-For-All Networks for Faster Multi-Platform DeploymentManas Sahni, Shreya Varshini, Alind Khare, Alexey TumanovICLR 2021 · 被引用 37 次
