Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
Junhan Kim, Chungman Lee, Eulrang Cho, Kyungphil Park, Ho-Young Kim, Joonyoung Kim, Yongkweon Jeon
Abstract
With the increasing complexity of generative AI models, post-training quantization (PTQ) has emerged as a promising solution for deploying hyper-scale models on edge devices such as mobile and TVs. Existing PTQ schemes, however, consume considerable time and resources, which could be a bottleneck in real situations where frequent model updates and multiple hyperparameter tunings are required. As a cost-effective alternative, learning-free PTQ schemes have been proposed. However, the performance is somewhat limited because they cannot consider the inter-layer dependency within the attention module, which is a significant feature of Transformers. In this paper, we thus propose a novel PTQ algorithm that balances accuracy and efficiency. The key idea of the proposed algorithm called aespa is to perform quantization layer-wise for efficiency while targeting attention-wise reconstruction to consider the cross-layer dependency. Through extensive experiments on various language models and complexity analysis, we demonstrate that aespa is accurate and efficient in quantizing Transformer models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99f23b28-498b-47f6-8ee9-b014383bbe4eCited by top-tier papers5
- TurboBoA: Faster and Exact Attention-aware Quantization without BackpropagationJunhan Kim, Yeo Jeong Park, Seungwoo Son, Chungman Lee et al.ICLR 2026 · 3 citations
- Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models QuantizationChenwei Jia, Baoting Li, Xuchong Zhang, Mingzhuo Wei et al.CVPR 2026 · 3 citations
- BoA: Attention-aware Post-training Quantization without BackpropagationJunhan Kim, Ho-Young Kim, Eulrang Cho, Chungman Lee et al.ICML 2025
- STLA: Spatiotemporal Lookahead Alignment for Post-Training QuantizationZuqi Zhang, Chenghe Sun, Xiangyi Chu, Wei-Han Yu et al.ICML 2026
- LogART: Pushing the Limit of Efficient Logarithmic Post-Training QuantizationJiawei Xu, Yi Zheng, Chenghe Sun, Taiyu Zhou et al.ICLR 2026
Builds on19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
Related papers
- RL-PTQ: RL-based Mixed Precision Quantization for Hybrid Vision TransformersEunji Kwon, Minxuan Zhou, Weihong Xu, Tajana Rosing et al.DAC 2024 · 4 citations
- CBQ: Cross-Block Quantization for Large Language ModelsXin Ding, Xiaoyu Liu, Zhijun Tu, Yun Zhang et al.ICLR 2025 · 1 citation
- EasyQuant: An Efficient Data-free Quantization Algorithm for LLMsHanlin Tang, Yifu Sun, Decheng Wu, Kai Liu et al.EMNLP 2023 · 4 citations
- GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric CalibrationYuhang Li, Ruokai Yin, Donghyun Lee, Shiting Xiao et al.ICML 2025
- DVD-Quant: Data-free Video Diffusion Transformers QuantizationZhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu et al.ICLR 2026 · 13 citations
