Any-Precision Deep Neural Networks
Haichao Yu, Haoxiang Li, Humphrey Shi, Thomas S. Huang, Gang Hua
Abstract
We present Any-Precision Deep Neural Networks (Any- Precision DNNs), which are trained with a new method that empowers learned DNNs to be flexible in any numerical precision during inference. The same model in runtime can be flexibly and directly set to different bit-width, by trun- cating the least significant bits, to support dynamic speed and accuracy trade-off. When all layers are set to low- bits, we show that the model achieved accuracy compara- ble to dedicated models trained at the same precision. This nice property facilitates flexible deployment of deep learn- ing models in real-world applications, where in practice trade-offs between model accuracy and runtime efficiency are often sought. Previous literature presents solutions to train models at each individual fixed efficiency/accuracy trade-off point. But how to produce a model flexible in runtime precision is largely unexplored. When the demand of efficiency/accuracy trade-off varies from time to time or even dynamically changes in runtime, it is infeasible to re-train models accordingly, and the storage budget may forbid keeping multiple models. Our proposed framework achieves this flexibility without performance degradation. More importantly, we demonstrate that this achievement is agnostic to model architectures. We experimentally validated our method with different deep network backbones (AlexNet-small, Resnet-20, Resnet-50) on different datasets (SVHN, Cifar-10, ImageNet) and observed consistent results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 326944f5-547d-4a9f-b8b8-b3c9e90cdd33Cited by top-tier papers20
- IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationYunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu et al.CVPR 2022 · 79 citations
- FFNeRV: Flow-Guided Frame-Wise Neural Representations for VideosJoo Chan Lee, Daniel Rho, Jong Hwan Ko, Eunbyung ParkACM MM 2023 · 59 citations
- Dynamic Network Quantization for Efficient Video InferenceXimeng Sun, Rameswar Panda, Chun-Fu (Richard) Chen, Aude Oliva et al.ICCV 2021 · 56 citations
- Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMsYeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim et al.ICML 2024 · 51 citations
- BatchQuant: Quantized-for-all Architecture Search with Robust QuantizerHaoping Bai, Meng Cao, Ping Huang, Jiulong ShanNeurIPS 2021 · 43 citations
Builds on5
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 725 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
Related papers
- InstantNet: Automated Generation and Deployment of Instantaneously Switchable-Precision NetworksYonggan Fu, Zhongzhi Yu, Yongan Zhang, Yifan Jiang et al.DAC 2021 · 6 citations
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural NetworksLéopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Oguz H. Elibol et al.ICLR 2020 · 53 citations
- Not All Bits have Equal Value: Heterogeneous Precisions via Trainable NoisePedro Savarese, Xin Yuan, Yanjing Li, Michael MaireNeurIPS 2022 · 9 citations
- Arbitrary Bit-width Network: A Joint Layer-Wise Quantization and Adaptive Inference ApproachChen Tang, Haoyu Zhai, Kai Ouyang, Zhi Wang et al.ACM MM 2022 · 15 citations
- AdaBits: Neural Network Quantization With Adaptive Bit-WidthsQing Jin, Linjie Yang, Zhenyu LiaoCVPR 2020
