FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
Jung Hyun Lee, Jeonghoon Kim, Se Jung Kwon, Dongsoo Lee
Abstract
Post-training quantization (PTQ) has been gaining popularity for the deployment of deep neural networks on resource-limited devices since unlike quantization-aware training, neither a full training dataset nor end-to-end training is required at all. As PTQ schemes based on reconstructing each layer or block output turn out to be effective to enhance quantized model performance, recent works have developed algorithms to devise and learn a new weight-rounding scheme so as to better reconstruct each layer or block output. In this work, we propose a simple yet effective new weight-rounding mechanism for PTQ, coined FlexRound, based on element-wise division instead of typical element-wise addition such that FlexRound enables jointly learning a common quantization grid size as well as a different scale for each pre-trained weight. Thanks to the reciprocal rule of derivatives induced by element-wise division, FlexRound is inherently able to exploit pre-trained weights when updating their corresponding scales, and thus, flexibly quantize pre-trained weights depending on their magnitudes. We empirically validate the efficacy of FlexRound on a wide range of models and tasks. To the best of our knowledge, our work is the first to carry out comprehensive experiments on not only image classification and natural language understanding but also natural language generation. Moreover, we demonstrate, for the first time, that large language models can be efficiently quantized, with only a negligible impact on performance compared to half-precision baselines, achieved by reconstructing the output in a block-by-block manner. Our code is available at https://github.com/onliwad101/FlexRound_LRQ.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer QuantizationJeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park et al.NeurIPS 2023 · 157 citations
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou et al.NeurIPS 2024 · 47 citations
- LoQT: Low-Rank Adapters for Quantized PretrainingSebastian Loeschcke, Mads Toftrup, Michael J. Kastoryano, Serge J. Belongie et al.NeurIPS 2024 · 14 citations
- PTMQ: Post-training Multi-Bit Quantization of Neural NetworksKe Xu, Zhongcheng Li, Shanshan Wang, Xingyi ZhangAAAI 2024 · 12 citations
- FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up TablesGunho Park, Hyeokjun Kwon, Jiwoo Kim, Jeongin Bae et al.HPCA 2025 · 10 citations
Builds on11
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
- Accurate Post Training Quantization With Small Calibration SetsItay Hubara, Yury Nahshan, Yair Hanani, Ron Banner et al.ICML 2021 · 238 citations
Related papers
- MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationChun Hu, Junhui He, Shangyu Wu, Yuxin He et al.EMNLP 2025 · 1 citation
- SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language ModelsHan Liu, Haotian Gao, Xiaotong Zhang, Changya Li et al.KDD 2025 · 1 citation
- SliderQuant: Accurate Post-Training Quantization for LLMsShigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan et al.ICLR 2026 · 6 citations
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang et al.NeurIPS 2022 · 62 citations
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionXi Zhang, Xiaolin Wu, Jiamang Wang, Weisi LinNeurIPS 2025 · 4 citations
