Lune

NeurIPS2025Top-tier venue

FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic

Kanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon, Jinho Lee

2025Year
1Citations

Abstract

Low-bit floating-point (FP) formats, such as FP8, provide significant acceleration and memory savings in model training thanks to native hardware support on modern GPUs and NPUs. However, we analyze that FP8 quantization offers speedup primarily for large-dimensional matrix multiplications, while inherent quantization overheads diminish speedup when applied to low-rank adaptation (LoRA), which uses small-dimensional matrices for efficient fine-tuning of large language models (LLMs). To address this limitation, we propose FALQON, a novel framework that eliminates the quantization overhead from separate LoRA computational paths by directly merging LoRA adapters into an FP8-quantized backbone during finetuning. Furthermore, we reformulate the forward and backward computations for merged adapters to significantly reduce quantization overhead, and introduce a row-wise proxy update mechanism that efficiently integrates substantial updates into the quantized backbone. Experimental evaluations demonstrate that FALQON achieves approximately a 3× training speedup over existing quantized LoRA methods with a similar level of accuracy, providing a practical solution for efficient large-scale model fine-tuning. Moreover, FALQON's end-to-end FP8 workflow removes the need for post-training quantization, facilitating efficient deployment. Code is available at https://github.com/iamkanghyunchoi/falqon.

These quantization overheads are especially critical when it comes to fine-tuning with low-rank adaptation (LoRA) [20]. LoRA inserts small-dimensional trainable low-rank matrices (adapters) to capture task-specific knowledge, significantly reducing the memory cost by using fewer trainable parameters. However, for matrices with small dimensions, such as LoRA adapters, the overhead incurred by FP8 quantization can outweigh the benefits from FP8 multiplications. Also, separate forward and backward paths for LoRA introduce a larger number of quantization operations, worsening the overall overhead. In our preliminary analyses (Section 4), we show that applying FP8 quantization to LoRA introduces significant quantization overhead, limiting speedup. This slowdown in LoRA poses critical challenges in practical scenarios, where numerous adapters must be trained to support personalization [59], multi-task learning [28], and rapid updates in dynamic, user-specific environments (see Section 3.1). Thus, efficient acceleration of LoRA fine-tuning is essential to enable timely, scalable, and cost-effective deployment of LLMs under practical computational constraints.

To address this, we propose FALQON (FP8-Accelerated LoRA Quantization), a novel framework designed specifically to accelerate FP8-based quantized LoRA fine-tuning by reducing quantization overheads. Instead of separate LoRA adapters, FALQON merges adapters directly into the FP8 backbone during fine-tuning, leveraging the initial quantization error as an implicit LoRA initialization (melded LoRA) to eliminate extra quantization steps. Additionally, we reformulate both forward and backward computational paths for efficient gradient calculation of the merged adapters. A row-wise proxy update mechanism selectively applies substantial weight updates to the backbone, avoiding ineffective updates that vanish under low-bit quantization and further enhancing overall efficiency.

Through extensive experiments on various tasks, we demonstrate that FALQON achieves up to 3× faster fine-tuning compared to quantized LoRA baselines, while maintaining comparable accuracy. Moreover, the end-to-end FP8 workflow of FALQON eliminates the need for post-training quantization, facilitating efficient deployment. Our key contributions are summarized as follows:

• We analyze FP8 quantization overhead and show that existing FP8 quantization methods primarily target large-dimensional matrix multiplications, resulting in substantial overhead and limited speedups when directly applied to LoRA's small-dimensional adapters.

• We propose FALQON, a novel framework that merges LoRA adapters into an FP8-quantized backbone during fine-tuning, significantly reducing quantization overhead.

• We reformulate forward and backward paths for efficient gradient computation of merged adapters and introduce a row-wise proxy update mechanism that selectively integrates substantial updates, avoiding unnecessary weight modifications under low-bit quantization.

• We empirically demonstrate that FALQON achieves up to 3× faster fine-tuning compared to existing methods, providing detailed breakdown analyses to identify sources of acceleration, while maintaining comparable accuracy across comprehensive evaluations.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cf667ce5-0f0e-4226-bfd6-446b10a5f7b7

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines