Arbitrary Bit-width Network: A Joint Layer-Wise Quantization and Adaptive Inference Approach
Chen Tang, Haoyu Zhai, Kai Ouyang, Zhi Wang, Yifei Zhu, Wenwu Zhu
Abstract
Conventional model quantization methods use a fixed quantization scheme to different data samples, which ignores the inherent"recognition difficulty" differences between various samples. We propose to feed different data samples with varying quantization schemes to achieve a data-dependent dynamic inference, at a fine-grained layer level. However, enabling this adaptive inference with changeable layer-wise quantization schemes is challenging because the combination of bit-widths and layers is growing exponentially, making it extremely difficult to train a single model in such a vast searching space and use it in practice. To solve this problem, we present the Arbitrary Bit-width Network (ABN), where the bit-widths of a single deep network can change at runtime for different data samples, with a layer-wise granularity. Specifically, first we build a weight-shared layer-wise quantizable "super-network" in which each layer can be allocated with multiple bit-widths and thus quantized differently on demand. The super-network provides a considerably large number of combinations of bit-widths and layers, each of which can be used during inference without retraining or storing myriad models. Second, based on the well-trained super-network, each layer's runtime bit-width selection decision is modeled as a Markov Decision Process (MDP) and solved by an adaptive inference strategy accordingly. Experiments show that the super-network can be built without accuracy degradation, and the bit-widths allocation of each layer can be adjusted to deal with various inputs on the fly. On ImageNet classification, we achieve 1.1% top1 accuracy improvement while saving 36.2% BitOps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd25e566-8bb4-474f-9fed-8e77b40a6818Cited by top-tier papers7
- NetLLM: Adapting Large Language Models for NetworkingDuo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang et al.SIGCOMM 2024 · 162 citations
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model AccelerationYe Li, Yuan Meng, Zewen Sun, Kangye Ji et al.ICLR 2026 · 60 citations
- ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile DevicesChen Tang, Li Lyna Zhang, Huiqiang Jiang, Jiahang Xu et al.ICCV 2023 · 15 citations
- SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin RegularizationChen Tang, Kai Ouyang, Zenghao Chai, Yunpeng Bai et al.ACM MM 2023 · 11 citations
- Retraining-free Model Quantization via One-Shot Weight-Coupling LearningChen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie et al.CVPR 2024 · 5 citations
Builds on14
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
Related papers
- Fractional Skipping: Towards Finer-Grained Dynamic CNN InferenceJianghao Shen, Yue Wang, Pengfei Xu, Yonggan Fu et al.AAAI 2020 · 49 citations
- Any-Precision Deep Neural NetworksHaichao Yu, Haoxiang Li, Humphrey Shi, Thomas S. Huang et al.AAAI 2021 · 79 citations
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma et al.CVPR 2022 · 38 citations
- AdaBM: On-the-Fly Adaptive Bit Mapping for Image Super-ResolutionCheeun Hong, Kyoung Mu LeeCVPR 2024
- AdaBits: Neural Network Quantization With Adaptive Bit-WidthsQing Jin, Linjie Yang, Zhenyu LiaoCVPR 2020
