ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization
Cong Guo, Chen Zhang, Jingwen Leng, Zihan Liu, Fan Yang, Yunxin Liu, Minyi Guo, Yuhao Zhu
Abstract
Quantization is a technique to reduce the computation and memory cost of DNN models, which are getting increasingly large. Existing quantization solutions use fixed-point integer or floating-point types, which have limited benefits, as both require more bits to maintain the accuracy of original models. On the other hand, variable-length quantization uses low-bit quantization for normal values and high-precision for a fraction of outlier values. Even though this line of work brings algorithmic benefits, it also introduces significant hardware overheads due to variable-length encoding and decoding.In this work, we propose a fixed-length a daptive n umerical data t ype called ANT to achieve low-bit quantization with tiny hardware overheads. Our data type ANT leverages two key innovations to exploit the intra-tensor and inter-tensor adaptive opportunities in DNN models. First, we propose a particular data type, flint, that combines the advantages of float and int for adapting to the importance of different values within a tensor. Second, we propose an adaptive framework that selects the best type for each tensor according to its distribution characteristics. We design a unified processing element architecture for ANT and show its ease of integration with existing DNN accelerators. Our design results in speedup and energy efficiency improvement over the state-of-the-art quantization accelerators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4890f2b3-ba24-4d2d-ba4e-e563f2f09cc3Cited by top-tier papers35
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationCong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng et al.ISCA 2023 · 151 citations
- Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime RequantizationJungi Lee, Wonbeom Lee, Jaewoong SimISCA 2024 · 41 citations
- Nimbus: Secure and Efficient Two-Party Inference for TransformersZhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu et al.NeurIPS 2024 · 34 citations
- GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory StitchingCong Guo, Rui Zhang, Jiale Xu, Jingwen Leng et al.ASPLOS 2024 · 30 citations
- LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupXiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang et al.MobiCom 2023 · 29 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
Related papers
- INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair EncodingFangxin Liu, Ning Yang, Zhiyan Song, Zongwu Wang et al.DAC 2024 · 10 citations
- M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical TypeWeiming Hu, Haoyan Zhang, Cong Guo, Yu Feng et al.HPCA 2025 · 19 citations
- Algorithm-Hardware Co-Design of Adaptive Floating-Point Encodings for Resilient Deep Learning InferenceThierry Tambe, En-Yu Yang, Zishen Wan, Yuntian Deng et al.DAC 2020 · 67 citations
- Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationLian Liu, Zhaohui Xu, Yintao He, Ying Wang et al.DAC 2024 · 5 citations
- BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationYuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang et al.HPCA 2025 · 23 citations
