FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
Kanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon, Jinho Lee
Abstract
Low-bit floating-point (FP) formats, such as FP8, provide significant acceleration and memory savings in model training thanks to native hardware support on modern GPUs and NPUs. However, we analyze that FP8 quantization offers speedup primarily for large-dimensional matrix multiplications, while inherent quantization overheads diminish speedup when applied to low-rank adaptation (LoRA), which uses small-dimensional matrices for efficient fine-tuning of large language models (LLMs). To address this limitation, we propose FALQON, a novel framework that eliminates the quantization overhead from separate LoRA computational paths by directly merging LoRA adapters into an FP8-quantized backbone during finetuning. Furthermore, we reformulate the forward and backward computations for merged adapters to significantly reduce quantization overhead, and introduce a row-wise proxy update mechanism that efficiently integrates substantial updates into the quantized backbone. Experimental evaluations demonstrate that FALQON achieves approximately a 3× training speedup over existing quantized LoRA methods with a similar level of accuracy, providing a practical solution for efficient large-scale model fine-tuning. Moreover, FALQON's end-to-end FP8 workflow removes the need for post-training quantization, facilitating efficient deployment. Code is available at https://github.com/iamkanghyunchoi/falqon.
These quantization overheads are especially critical when it comes to fine-tuning with low-rank adaptation (LoRA) [20]. LoRA inserts small-dimensional trainable low-rank matrices (adapters) to capture task-specific knowledge, significantly reducing the memory cost by using fewer trainable parameters. However, for matrices with small dimensions, such as LoRA adapters, the overhead incurred by FP8 quantization can outweigh the benefits from FP8 multiplications. Also, separate forward and backward paths for LoRA introduce a larger number of quantization operations, worsening the overall overhead. In our preliminary analyses (Section 4), we show that applying FP8 quantization to LoRA introduces significant quantization overhead, limiting speedup. This slowdown in LoRA poses critical challenges in practical scenarios, where numerous adapters must be trained to support personalization [59], multi-task learning [28], and rapid updates in dynamic, user-specific environments (see Section 3.1). Thus, efficient acceleration of LoRA fine-tuning is essential to enable timely, scalable, and cost-effective deployment of LLMs under practical computational constraints.
To address this, we propose FALQON (FP8-Accelerated LoRA Quantization), a novel framework designed specifically to accelerate FP8-based quantized LoRA fine-tuning by reducing quantization overheads. Instead of separate LoRA adapters, FALQON merges adapters directly into the FP8 backbone during fine-tuning, leveraging the initial quantization error as an implicit LoRA initialization (melded LoRA) to eliminate extra quantization steps. Additionally, we reformulate both forward and backward computational paths for efficient gradient calculation of the merged adapters. A row-wise proxy update mechanism selectively applies substantial weight updates to the backbone, avoiding ineffective updates that vanish under low-bit quantization and further enhancing overall efficiency.
Through extensive experiments on various tasks, we demonstrate that FALQON achieves up to 3× faster fine-tuning compared to quantized LoRA baselines, while maintaining comparable accuracy. Moreover, the end-to-end FP8 workflow of FALQON eliminates the need for post-training quantization, facilitating efficient deployment. Our key contributions are summarized as follows:
• We analyze FP8 quantization overhead and show that existing FP8 quantization methods primarily target large-dimensional matrix multiplications, resulting in substantial overhead and limited speedups when directly applied to LoRA's small-dimensional adapters.
• We propose FALQON, a novel framework that merges LoRA adapters into an FP8-quantized backbone during fine-tuning, significantly reducing quantization overhead.
• We reformulate forward and backward paths for efficient gradient computation of merged adapters and introduce a row-wise proxy update mechanism that selectively integrates substantial updates, avoiding unnecessary weight modifications under low-bit quantization.
• We empirically demonstrate that FALQON achieves up to 3× faster fine-tuning compared to existing methods, providing detailed breakdown analyses to identify sources of acceleration, while maintaining comparable accuracy across comprehensive evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf667ce5-0f0e-4226-bfd6-446b10a5f7b7Builds on26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and ServingYuchen Zhang, Hanyue Du, Chun Cao, Jingwei XuNeurIPS 2025 · 1 citation
- IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion ModelsHang Guo, Yawei Li, Tao Dai, Shu-Tao Xia et al.ICML 2025
- LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model FinetuningHan Guo, Philip Greengard, Eric P. Xing, Yoon KimICLR 2024 · 94 citations
- L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language ModelsHyesung Jeon, Yulhwa Kim, Jae-Joon KimACL 2025
- ProjQ: Project-and-Quantize for Adapter-Aware LLM CompressionWenya Yu, Chao Zhang, Li Wang, Samson Lasaulce et al.ICML 2026
