Lune

INFOCOM2025顶会

Federated Adaptive Fine-Tuning of Large Language Models with Heterogeneous Quantization and LoRA

Zhidong Gao, Zhenxiao Zhang, Yuanxiong Guo, Yanmin Gong

2025年份
12被引次数
5顶会引用

摘要

Federated learning (FL) with parameter-efficient fine-tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), offers a privacy-preserving solution for fine-tuning large language models (LLMs) on edge devices. However, due to the enormous size of LLMs, federated fine-tuning with LoRA still faces significant challenges, including high training latency and substantial memory requirements. To address these challenges, we propose FAH-QLoRA, a novel time- and memory-efficient federated adaptive fine-tuning framework for LLMs, which leverages heterogeneous model quantization and LoRA. The key idea behind FAH-QLoRA is to quantize the base model within LoRA and dynamically adjust the LoRA ranks, enabling efficient fine-tuning on resource-constrained and heterogeneous devices while reducing training latency and maintaining performance. FAH-QLoRA integrates two key techniques: i) dynamically adjusting LoRA ranks across training rounds to promote faster convergence and lower resource usage, and ii) assigning heterogeneous LoRA ranks across devices to mitigate the straggler effect during the FL process, ensuring that heterogeneous resource constraints are met. We analyze the convergence of FAH-QLoRA under general non-convex and non-IID FL settings. Extensive experiments demonstrate that FAH-QLoRA can reduce training time by up to 45.86% and memory usage by up to 44.15% compared to baseline methods.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖