BatchQuant: Quantized-for-all Architecture Search with Robust Quantizer
Haoping Bai, Meng Cao, Ping Huang, Jiulong Shan
Abstract
As the applications of deep learning models on edge devices increase at an accelerating pace, fast adaptation to various scenarios with varying resource constraints has become a crucial aspect of model deployment. As a result, model optimization strategies with adaptive configuration are becoming increasingly popular. While single-shot quantized neural architecture search enjoys flexibility in both model architecture and quantization policy, the combined search space comes with many challenges, including instability when training the weight-sharing supernet and difficulty in navigating the exponentially growing search space. Existing methods tend to either limit the architecture search space to a small set of options or limit the quantization policy search space to fixed precision policies. To this end, we propose BatchQuant, a robust quantizer formulation that allows fast and stable training of a compact, single-shot, mixed-precision, weight-sharing supernet. We employ BatchQuant to train a compact supernet (offering over quantized subnets) within substantially fewer GPU hours than previous methods. Our approach, Quantized-for-all (QFA), is the first to seamlessly extend one-shot weight-sharing NAS supernet to support subnets with arbitrary ultra-low bitwidth mixed-precision quantization policies without retraining. QFA opens up new possibilities in joint hardware-aware neural architecture search and quantization. We demonstrate the effectiveness of our method on ImageNet and achieve SOTA Top-1 accuracy under a low complexity constraint ( MFLOPs). The code and models will be made publicly available at https://github.com/bhpfelix/QFA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa9cc572-8f5d-45a3-9374-9f88b586a078Cited by top-tier papers6
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu et al.NeurIPS 2023 · 24 citations
- MixPath: A Unified Approach for One-shot Neural Architecture SearchXiangxiang Chu, Shun Lu, Xudong Li, Bo ZhangICCV 2023 · 24 citations
- EQ-Net: Elastic Quantization Neural NetworksKe Xu, Lei Han, Ye Tian, Shangshang Yang et al.ICCV 2023 · 21 citations
- ClimbQ: Class Imbalanced Quantization Enabling Robustness on Efficient InferencesTing-An Chen, De-Nian Yang, Ming-Syan ChenNeurIPS 2022 · 8 citations
- JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-ExplorationMingzi Wang, Yuan Meng, Chen Tang, Weixiang Zhang et al.AAAI 2025 · 3 citations
Builds on12
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad et al.NeurIPS 2020 · 149 citations
Related papers
- No Retraining at Edge: Efficient Resource-Aware Mixed-Precision Quantization via Federated Supernet LearningLianbo Ma, Yonghui Su, Nan Li, Xingwei WangICML 2026
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang et al.CVPR 2022 · 9 citations
- Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture SearchMingzhu Shen, Feng Liang, Ruihao Gong, Yuhang Li et al.ICCV 2021 · 50 citations
- Double Rounding: Nearly Lossless Adaptive Bit Switching for QATHaiduo Huang, Zhenhua Liu, Tian Xia, Pengju RenAAAI 2026
- One-Shot Model for Mixed-Precision QuantizationIvan Koryakovskiy, Alexandra Yakovleva, Valentin Buchnev, Temur Isaev et al.CVPR 2023
