BoA: Attention-aware Post-training Quantization without Backpropagation
Junhan Kim, Ho-Young Kim, Eulrang Cho, Chungman Lee, Joonyoung Kim, Yongkweon Jeon
Abstract
Post-training quantization (PTQ) is a promising solution for deploying large language models (LLMs) on resource-constrained devices. Early methods developed for small-scale networks, such as ResNet, rely on gradient-based optimization, which becomes impractical for hyper-scale LLMs with billions of parameters. While recently proposed backpropagation-free or transformation-based methods alleviate this issue, they ignore inter-layer interactions or use the naive nearest-rounding-based quantized weight assignment to save the heavy computational cost of weight optimization. In this paper, we introduce a novel backpropagation-free PTQ algorithm that optimizes quantized weights by considering inter-layer dependencies. The key innovation is the development of attention-aware Hessian matrices that capture inter-layer interactions within the attention module. Extensive experiments demonstrate that our approach not only outperforms existing weight quantization methods but also shows good synergy with conventional methods to suppress activation outliers, leading to state-of-the-art weight-activation quantization performance. The code will be available at https://github.com/SamsungLabs/BoA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eec39a3-a6ea-462c-b855-e8e2d78596dbCited by top-tier papers2
- Block Rotation is All You Need for MXFP4 QuantizationYuantian Shao, Peisong Wang, Yuanteng Chen, Chang Xu et al.ICML 2026 · 16 citations
- TurboBoA: Faster and Exact Attention-aware Quantization without BackpropagationJunhan Kim, Yeo Jeong Park, Seungwoo Son, Chungman Lee et al.ICLR 2026 · 3 citations
Builds on15
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 440 citations
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
Related papers
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language ModelsCuong Pham, Dung Anh Hoang, Cuong C. Nguyen, Trung Le et al.ACL 2026
- ACBQ: Adaptive Cross-Block Quantization of Large Language ModelsHailing Wang, Jianglin Lu, Yitian Zhang, Huimin Zeng et al.ACL 2026
- GuidedQuant: Large Language Model Quantization via Exploiting End Loss GuidanceJinuk Kim, Marwa El Halabi, Wonpyo Park, Clemens J. S. Schaefer et al.ICML 2025
- OAC: Output-adaptive Calibration for Accurate Post-training QuantizationAli Edalati, Alireza Ghaffari, Mahsa Ghazvini Nejad, Lu Hou et al.AAAI 2025 · 8 citations
- Towards Next-Level Post-Training Quantization of Hyper-Scale TransformersJunhan Kim, Chungman Lee, Eulrang Cho, Kyungphil Park et al.NeurIPS 2024 · 10 citations
