APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
Tianzhe Wang, Kuan Wang, Han Cai, Ji Lin, Zhijian Liu, Hanrui Wang, Yujun Lin, Song Han
Abstract
We present APQ, a novel design methodology for efficient deep learning deployment. Unlike previous methods that separately optimize the neural network architecture, pruning policy, and quantization policy, we design to optimize them in a joint manner. To deal with the larger design space it brings, we devise to train a quantizationaware accuracy predictor that is fed to the evolutionary search to select the best fit. Since directly training such a predictor requires time-consuming quantization data collection, we propose to use predictor-transfer technique to get the quantization-aware predictor: we first generate a large dataset of NN architecture, ImageNet accuracy pairs by sampling a pretrained unified once-for-all network and doing direct evaluation; then we use these data to train an accuracy predictor without quantization, followed by transferring its weights to train the quantization-aware predictor, which largely reduces the quantization data collection time. Extensive experiments on ImageNet show the benefits of this joint design methodology: the model searched by our method maintains the same level accuracy as ResNet34 8-bit model while saving 8× BitOps; we achieve 2×/1.3× latency/energy saving compared to 36] while obtaining the same level accuracy; the marginal search cost of joint optimization for a new deployment scenario outperforms separate optimizations using ProxylessNAS+AMC+HAQ [5, 12, 36] by 2.3% accuracy while reducing orders of magnitude GPU hours and CO 2 emission with respect to the training cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- EfficientViT: Lightweight Multi-Scale Attention for High-Resolution Dense PredictionHan Cai, Junyan Li, Muyan Hu, Chuang Gan et al.ICCV 2023 · 265 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- QuantumNAS: Noise-Adaptive Search for Robust Quantum CircuitsHanrui Wang, Yongshan Ding, Jiaqi Gu, Yujun Lin et al.HPCA 2022 · 199 citations
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
Builds on3
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
Related papers
- SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 InferenceXudong Wang, Li Lyna Zhang, Jiahang Xu, Quanlu Zhang et al.ICCV 2023 · 3 citations
- Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture SearchMingzhu Shen, Feng Liang, Ruihao Gong, Yuhang Li et al.ICCV 2021 · 50 citations
- JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-ExplorationMingzi Wang, Yuan Meng, Chen Tang, Weixiang Zhang et al.AAAI 2025 · 3 citations
- Hardware-adaptive Efficient Latency Prediction for NAS via Meta-LearningHayeon Lee, Sewoong Lee, Song Chong, Sung Ju HwangNeurIPS 2021 · 32 citations
- CompOFA - Compound Once-For-All Networks for Faster Multi-Platform DeploymentManas Sahni, Shreya Varshini, Alind Khare, Alexey TumanovICLR 2021 · 37 citations
