AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-Tuning
Changhai Zhou, Shiyang Zhang, Yuhua Zhou, Qian Qiao, Jun Gao, Cheng Jin, KAIZHOU QIN, Weizhong Zhang
摘要
Quantization followed by parameter-efficient fine-tuning has emerged as a promising paradigm for downstream adaptation under tight GPU memory constraints. However, this sequential pipeline fails to leverage the intricate interaction between quantization bit-width and LoRA rank. Specifically, a carefully optimized quantization allocation with low quantization error does not always translate to strong fine-tuning performance, and different bit-width and rank configurations can lead to significantly varying outcomes under the same memory budget. To address this limitation, we propose AutoQRA, a joint optimization framework that simultaneously optimizes the bit-width and LoRA rank configuration for each layer during the mixed quantized fine-tuning process. To tackle the challenges posed by the large discrete search space and the high evaluation cost associated with frequent fine-tuning iterations, AutoQRA decomposes the optimization process into two stages. First, it first conducts a global multi-fidelity evolutionary search, where the initial population is warm-started by injecting layer-wise importance priors. This stage employs specific operators and a performance model to efficiently screen candidate configurations. Second, trust-region Bayesian optimization is applied to locally refine promising regions of the search space and identify optimal configurations under the given memory budget. This approach enables active compensation for quantization noise in specific layers during training. Experiments show that AutoQRA achieves performance close to full-precision fine-tuning with a memory footprint comparable to uniform 4-bit methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language ModelsYixiao Li, Yifan Yu, Chen Liang, Nikos Karampatziakis 等ICLR 2024 · 被引用 217 次
- QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language ModelsYuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen 等ICLR 2024 · 被引用 179 次
- Bayesian Optimisation over Multiple Continuous and Categorical InputsBin Xin Ru, Ahsan S. Alvi, Vu Nguyen, Michael A. Osborne 等ICML 2020 · 被引用 119 次
相关 Paper
- LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 BitsZikai Zhou, Qizheng Zhang, Hermann Kumbong, Kunle OlukotunICML 2025
- IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion ModelsHang Guo, Yawei Li, Tao Dai, Shu-Tao Xia 等ICML 2025
- Flat-LoRA: Low-Rank Adaptation over a Flat Loss LandscapeTao Li, Zhengbao He, Yujun Li, Yasheng Wang 等ICML 2025
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMsRongguang Ye, Ming Tang, Edith NgaiICLR 2026 · 被引用 1 次
- L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language ModelsHyesung Jeon, Yulhwa Kim, Jae-Joon KimACL 2025
