QERA: an Analytical Framework for Quantization Error Reconstruction
Cheng Zhang, Jeffrey T. H. Wong, Can Xiao, George Anthony Constantinides, Yiren Zhao
Abstract
The growing number of parameters and computational demands of large language models (LLMs) present significant challenges for their efficient deployment. Recently, there is an increasing interest in quantizing weights to extremely low precision while offsetting the resulting error with low-rank, high-precision error reconstruction terms. The combination of quantization and low-rank approximation is now popular in both adapter-based, parameter-efficient fine-tuning methods such as LoftQ (Li et al., 2023) and low-precision inference techniques including ZeroQuant-V2 (Yao et al., 2023) . Usually, the low-rank terms are calculated via the singular value decomposition (SVD) of the weight quantization error, minimizing the Frobenius and spectral norms of the weight approximation error. Recent methods like LQ-LoRA (Guo et al., 2023) and LQER (Zhang et al., 2024a) introduced hand-crafted heuristics to minimize errors in layer outputs (activations) rather than weights, resulting improved quantization results. However, these heuristic-based methods lack an analytical solution to guide the design of quantization error reconstruction terms. In this paper, we revisit this problem and formulate an analytical framework, named Quantization Error Reconstruction Analysis (QERA), and offer a closed-form solution to the problem. We show QERA benefits both existing low-precision fine-tuning and inference methods -QERA achieves a fine-tuned accuracy gain for ∆ acc = 6.05% of 2-bit RoBERTabase on GLUE compared to LoftQ; and obtains ∆ acc = 2.97% higher post-training quantization accuracy of 4-bit Llama-3.1-70B compared to ZeroQuant-V2 and ∆ ppl = -0.28 lower perplexity on WikiText2 compared to LQER. We open-source our code and models at github.com/ChengZhang-98/QERA. Published as a conference paper at ICLR 2025 precision. Works such as ZeroQuant-V2 (Yao et al., 2023) and LQER (Zhang et al., 2024a) have shown that adding a high-precision low-rank component, as low as 8 or 32, can recover considerable model performance for 3-or 4-bit weight quantization. Although both the QPEFT and PTQ methods have demonstrated substantial performance improvements in lowering the computational overhead of LLMs, a theoretical analysis of quantization error reconstruction is lacking. Usually, A k and B k are calculated by applying truncated singular value decomposition (SVD) to the weight quantization error (W -W ), minimizing the Frobenius and spectral norms of the weight approximation error. However, recent work on activation-aware quantization and knowledge distillation implies that minimizing layer output error may lead to a greater performance gain than minimizing weight approximation error (Lin et al., 2024; Liu et al., 2023a; Shao et al., 2023) . Besides the unsettled minimization objective, it has remained unclear whether there exists a theoretically optimal solution for the values of A k and B k , and if so, how one can solve for it. A better initialization or theoretically grounded initialization of A k and B k brings direct benefits for both QPEFT and PTQ. In QPEFT, the initialization of LoRA (Hu et al., 2021) , which uses element-wise Gaussian random values for A k and zeros for B k , struggles under aggressive quantization since the quantization error can derail fine-tuning. In PTQ, the quantized model performance is based on the computation of the low-rank terms, given a specific quantization function q(•) and rank k. In this paper, we aim to provide an analytical framework for the quantization error reconstruction problem. To demonstrate the effectiveness of our theoretical framework, we further apply our analytical solutions to state-of-the-art QPEFT and PTQ methods and show the significant performance improvements under the same computational budget. Specifically, our contributions are as follows: • We show that the commonly used objective for solving the quantization error reconstruction problem in prior work , i.e., minimizing the weight approximation error (e.g., ||W -W || p ), does not guarantee a reduced model output error. Instead, we show that minimizing the layer output error (e.g., ||y -y|| p ) is closely related to minimizing the model output error. • We derive the analytical solution to the low-rank terms A k and B k by minimizing the layer output error. We demonstrate that under a statistical assumption, this solution can be found in a particularly computationally efficient manner, also explaining the success of LQER. • We empirically demonstrate the effectiveness of our solutions by applying them to stateof-the-art QPEFT and PTQ methods. Our analytical framework, QERA, significantly improves the performance of these methods. For example, QERA achieves ∆ acc = 6.05% higher accuracy of 2-bit RoBERTa-base on GLUE compared to LoftQ, improving the finetuning accuracy and efficiency. Moreover, QERA obtains ∆ acc = 2.97% higher accuracy than ZeroQuant-V2, when quantizing LLaMA-3-70B to 4 bits, averaged across six tasks. This
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4d86824-8d29-44d4-aa61-4bc5d4081392Cited by top-tier papers6
- A3: an Analytical Low-Rank Approximation Framework for AttentionJeffrey T. H. Wong, Cheng Zhang, Xinye Cao, Pedro Gimenes et al.ICML 2026 · 4 citations
- Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMsYoonjun Cho, Dongjae Jeon, Soeun Kim, Moongyu Jeon et al.ICML 2026 · 1 citation
- SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM QuantizationYeonsik Park, Hyeonseong Kim, Seungkyu ChoiICLR 2026 · 1 citation
- GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMsSelim An, Il hong Suh, Yeseong KimICLR 2026
- TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D TilingHongyaoxing Gu, Xinzhe Chen, LIJUAN HU, Liu fangfangICML 2026
Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- LQER: Low-Rank Quantization Error Reconstruction for LLMsCheng Zhang, Jianyi Cheng, George Anthony Constantinides, Yiren ZhaoICML 2024 · 33 citations
- ASER: Activation Smoothing and Error Reconstruction for Large Language Model QuantizationWeibo Zhao, Yubin Shi, Xinyu Lyu, Wanchen Sui et al.AAAI 2025 · 7 citations
- ProjQ: Project-and-Quantize for Adapter-Aware LLM CompressionWenya Yu, Chao Zhang, Li Wang, Samson Lasaulce et al.ICML 2026
- L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language ModelsHyesung Jeon, Yulhwa Kim, Jae-Joon KimACL 2025
- CBQ: Cross-Block Quantization for Large Language ModelsXin Ding, Xiaoyu Liu, Zhijun Tu, Yun Zhang et al.ICLR 2025 · 1 citation
