Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
Haotong Qin, Xudong Ma, Xingyu Zheng, Xiaoyang Li, Yang Zhang, Shouda Liu, Jie Luo, Xianglong Liu, Michele Magno
摘要
The LoRA-finetuning quantization of LLMs has been extensively studied to obtain accurate yet compact LLMs for deployment on resource-constrained hardware. However, existing methods cause the quantized LLM to severely degrade and even fail to benefit from the finetuning of LoRA. This paper proposes a novel IR-QLoRA for pushing quantized LLMs with LoRA to be highly accurate through information retention. The proposed IR-QLoRA mainly relies on two technologies derived from the perspective of unified information: (1) statistics-based Information Calibration Quantization allows the quantized parameters of LLM to retain original information accurately; (2) finetuning-based Information Elastic Connection makes LoRA utilizes elastic representation transformation with diverse information. Comprehensive experiments show that IR-QLoRA can significantly improve accuracy across LLaMA and LLaMA2 families under 2-4 bit-widths, e.g., 4- bit LLaMA-7B achieves 1.4% improvement on MMLU compared with the state-of-the-art methods. The significant performance gain requires only a tiny 0.31% additional time consumption, revealing the satisfactory efficiency of our IR-QLoRA. We highlight that IR-QLoRA enjoys excellent versatility, compatible with various frameworks (e.g., NormalFloat and Integer quantization) and brings general accuracy gains. The code is available at https://github.com/htqin/ir-qlora.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Image Fusion via Vision-Language ModelZixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui 等ICML 2024 · 被引用 79 次
- Compressing Large Language Models by Joint Sparsification and QuantizationJinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu 等ICML 2024 · 被引用 33 次
- Q-SNNs: Quantized Spiking Neural NetworksWenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao 等ACM MM 2024 · 被引用 23 次
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu 等ICLR 2026 · 被引用 20 次
- BiDM: Pushing the Limit of Quantization for Diffusion ModelsXingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma 等NeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
相关 Paper
- QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language ModelsYuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen 等ICLR 2024 · 被引用 179 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language ModelsYixiao Li, Yifan Yu, Chen Liang, Nikos Karampatziakis 等ICLR 2024 · 被引用 217 次
- ApiQ: Finetuning of 2-Bit Quantized Large Language ModelBaohao Liao, Christian Herold, Shahram Khadivi, Christof MonzEMNLP 2024 · 被引用 2 次
- LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model FinetuningHan Guo, Philip Greengard, Eric P. Xing, Yoon KimICLR 2024 · 被引用 94 次
