FlatQuant: Flatness Matters for LLM Quantization
Yuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao, Kang Zhao, Yuening Li, Jiaxin Hu, Xianzhi Yu, Lu Hou, Chun Yuan, Xin Jiang, Wulong Liu, Jun Yao
Abstract
Recently, quantization has been widely used for the compression and acceleration of large language models (LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with equally spaced quantization points. Prior research explores various pre-quantization transformations to suppress outliers, such as per-channel scaling and Hadamard transformation. However, we observe that these transformed weights and activations can still exhibit steep and dispersed distributions. In this paper, we propose FLATQUANT (Fast and Learnable Affine Transformation), a new posttraining quantization approach that enhances the flatness of weights and activations. Our approach identifies optimal affine transformations for each linear layer, calibrated in hours via a lightweight objective. To reduce runtime overhead of affine transformation, we apply Kronecker product with two lightweight matrices, and fuse all operations in FLATQUANT into a single kernel. Extensive experiments demonstrate that FLATQUANT establishes a new state-of-the-art benchmark for quantization. For example, it achieves less than 1% accuracy drop for W4A4 quantization on the LLaMA-3-70B model, surpassing SpinQuant by 7.5%. Additionally, it provides up to 2.3x prefill speedup and 1.7x decoding speedup compared to the FP16 model. Code is available at: https:// github.com/ruikangliu/FlatQuant .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d900c101-83fa-4f3f-8dae-69ec0f507b2eCited by top-tier papers31
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM InferenceYesheng Liang, Haisheng Chen, Song Han, Zhijian LiuICLR 2026 · 19 citations
- FPTQuant: Function-Preserving Transforms for LLM QuantizationBoris van Breugel, Yelysei Bondarenko, Paul Whatmough, Markus NagelICML 2026 · 13 citations
- LLM.265: Video Codecs are Secretly Tensor CodecsCeyu Xu, Yongji Wu, Xinyu Yang, Beidi Chen et al.MICRO 2025 · 13 citations
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language ModelsWenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang et al.ICLR 2026 · 12 citations
- Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM DeploymentDeokjae Lee, Hyun Oh SongNeurIPS 2025 · 9 citations
Builds on17
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
Related papers
- Theory-optimal Quantization Based on FlatnessXiusheng Huang, Zhe Li, Xuanwu Yin, Lu Wang et al.ACL 2026
- AffineQuant: Affine Transformation Quantization for Large Language ModelsYuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling et al.ICLR 2024 · 56 citations
- SpinQuant: LLM Quantization with Learned RotationsZechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran et al.ICLR 2025
- LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware GridTianyi Zhang, Anshumali ShrivastavaICLR 2025
- OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution FittingXing Hu, Yuan Cheng, Dawei Yang, Zhixuan Chen et al.ICLR 2025
