LQER: Low-Rank Quantization Error Reconstruction for LLMs
Cheng Zhang, Jianyi Cheng, George Anthony Constantinides, Yiren Zhao
Abstract
Post-training quantization of Large Language Models (LLMs) is challenging. In this work, we introduce Low-rank Quantization Error Reduction (LQER), which combines quantization and low-rank approximation to recover the model capability. LQER leverages an activation-induced scale matrix to drive the singular value distribution of quantization error towards a desirable distribution, which enables nearly-lossless W4A8 quantization on various LLMs and downstream tasks without the need for knowledge distillation, grid search, or gradient-base iterative optimization. Unlike existing methods, the computation pattern of LQER eliminates the need for specialized Scatter and Gather processes to collect high-precision weights from irregular memory locations. Our W4A8 LLMs achieve near-lossless performance on six popular downstream tasks, while using 1.36 fewer hardware resources than the leading state-of-the-art method. We open-source our framework at https://github.com/ChengZhang-98/lqer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a59613c9-2519-42e0-a535-5cf8e3182f2fCited by top-tier papers17
- Compressing Large Language Models using Low Rank and Low Precision DecompositionRajarshi Saha, Naomi Sagan, Varun Srivastava, Andrea Goldsmith et al.NeurIPS 2024 · 74 citations
- PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM CompressionVladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev et al.NeurIPS 2024 · 69 citations
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal ModelsPengcheng Zheng, Chaoning Zhang, Jiarong Mo, Guohui Li et al.ICLR 2026 · 9 citations
- RILQ: Rank-Insensitive LoRA-Based Quantization Error Compensation for Boosting 2-Bit Large Language Model AccuracyGeonho Lee, Janghwan Lee, Sukjin Hong, Minsoo Kim et al.AAAI 2025 · 7 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- Outlier Suppression: Pushing the Limit of Low-bit Transformer Language ModelsXiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong et al.NeurIPS 2022 · 238 citations
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 196 citations
- Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointBita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu et al.NeurIPS 2020 · 153 citations
Related papers
- ASER: Activation Smoothing and Error Reconstruction for Large Language Model QuantizationWeibo Zhao, Yubin Shi, Xinyu Lyu, Wanchen Sui et al.AAAI 2025 · 7 citations
- QERA: an Analytical Framework for Quantization Error ReconstructionCheng Zhang, Jeffrey T. H. Wong, Can Xiao, George Anthony Constantinides et al.ICLR 2025
- SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM QuantizationYeonsik Park, Hyeonseong Kim, Seungkyu ChoiICLR 2026 · 1 citation
- Theory-optimal Quantization Based on FlatnessXiusheng Huang, Zhe Li, Xuanwu Yin, Lu Wang et al.ACL 2026
- RUQuant: Towards Refining Uniform Quantization for Large Language ModelsHan Liu, Haotian Gao, Changya Li, Feng Zhang et al.KDD 2026
