FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
Kanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon, Jinho Lee
摘要
Low-bit floating-point (FP) formats, such as FP8, provide significant acceleration and memory savings in model training thanks to native hardware support on modern GPUs and NPUs. However, we analyze that FP8 quantization offers speedup primarily for large-dimensional matrix multiplications, while inherent quantization overheads diminish speedup when applied to low-rank adaptation (LoRA), which uses small-dimensional matrices for efficient fine-tuning of large language models (LLMs). To address this limitation, we propose FALQON, a novel framework that eliminates the quantization overhead from separate LoRA computational paths by directly merging LoRA adapters into an FP8-quantized backbone during finetuning. Furthermore, we reformulate the forward and backward computations for merged adapters to significantly reduce quantization overhead, and introduce a row-wise proxy update mechanism that efficiently integrates substantial updates into the quantized backbone. Experimental evaluations demonstrate that FALQON achieves approximately a 3× training speedup over existing quantized LoRA methods with a similar level of accuracy, providing a practical solution for efficient large-scale model fine-tuning. Moreover, FALQON's end-to-end FP8 workflow removes the need for post-training quantization, facilitating efficient deployment. Code is available at https://github.com/iamkanghyunchoi/falqon.
These quantization overheads are especially critical when it comes to fine-tuning with low-rank adaptation (LoRA) [20]. LoRA inserts small-dimensional trainable low-rank matrices (adapters) to capture task-specific knowledge, significantly reducing the memory cost by using fewer trainable parameters. However, for matrices with small dimensions, such as LoRA adapters, the overhead incurred by FP8 quantization can outweigh the benefits from FP8 multiplications. Also, separate forward and backward paths for LoRA introduce a larger number of quantization operations, worsening the overall overhead. In our preliminary analyses (Section 4), we show that applying FP8 quantization to LoRA introduces significant quantization overhead, limiting speedup. This slowdown in LoRA poses critical challenges in practical scenarios, where numerous adapters must be trained to support personalization [59], multi-task learning [28], and rapid updates in dynamic, user-specific environments (see Section 3.1). Thus, efficient acceleration of LoRA fine-tuning is essential to enable timely, scalable, and cost-effective deployment of LLMs under practical computational constraints.
To address this, we propose FALQON (FP8-Accelerated LoRA Quantization), a novel framework designed specifically to accelerate FP8-based quantized LoRA fine-tuning by reducing quantization overheads. Instead of separate LoRA adapters, FALQON merges adapters directly into the FP8 backbone during fine-tuning, leveraging the initial quantization error as an implicit LoRA initialization (melded LoRA) to eliminate extra quantization steps. Additionally, we reformulate both forward and backward computational paths for efficient gradient calculation of the merged adapters. A row-wise proxy update mechanism selectively applies substantial weight updates to the backbone, avoiding ineffective updates that vanish under low-bit quantization and further enhancing overall efficiency.
Through extensive experiments on various tasks, we demonstrate that FALQON achieves up to 3× faster fine-tuning compared to quantized LoRA baselines, while maintaining comparable accuracy. Moreover, the end-to-end FP8 workflow of FALQON eliminates the need for post-training quantization, facilitating efficient deployment. Our key contributions are summarized as follows:
• We analyze FP8 quantization overhead and show that existing FP8 quantization methods primarily target large-dimensional matrix multiplications, resulting in substantial overhead and limited speedups when directly applied to LoRA's small-dimensional adapters.
• We propose FALQON, a novel framework that merges LoRA adapters into an FP8-quantized backbone during fine-tuning, significantly reducing quantization overhead.
• We reformulate forward and backward paths for efficient gradient computation of merged adapters and introduce a row-wise proxy update mechanism that selectively integrates substantial updates, avoiding unnecessary weight modifications under low-bit quantization.
• We empirically demonstrate that FALQON achieves up to 3× faster fine-tuning compared to existing methods, providing detailed breakdown analyses to identify sources of acceleration, while maintaining comparable accuracy across comprehensive evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and ServingYuchen Zhang, Hanyue Du, Chun Cao, Jingwei XuNeurIPS 2025 · 被引用 1 次
- IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion ModelsHang Guo, Yawei Li, Tao Dai, Shu-Tao Xia 等ICML 2025
- LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model FinetuningHan Guo, Philip Greengard, Eric P. Xing, Yoon KimICLR 2024 · 被引用 94 次
- L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language ModelsHyesung Jeon, Yulhwa Kim, Jae-Joon KimACL 2025
- ProjQ: Project-and-Quantize for Adapter-Aware LLM CompressionWenya Yu, Chao Zhang, Li Wang, Samson Lasaulce 等ICML 2026
