Federated Adaptive Fine-Tuning of Large Language Models with Heterogeneous Quantization and LoRA
Zhidong Gao, Zhenxiao Zhang, Yuanxiong Guo, Yanmin Gong
Abstract
Federated learning (FL) with parameter-efficient fine-tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), offers a privacy-preserving solution for fine-tuning large language models (LLMs) on edge devices. However, due to the enormous size of LLMs, federated fine-tuning with LoRA still faces significant challenges, including high training latency and substantial memory requirements. To address these challenges, we propose FAH-QLoRA, a novel time- and memory-efficient federated adaptive fine-tuning framework for LLMs, which leverages heterogeneous model quantization and LoRA. The key idea behind FAH-QLoRA is to quantize the base model within LoRA and dynamically adjust the LoRA ranks, enabling efficient fine-tuning on resource-constrained and heterogeneous devices while reducing training latency and maintaining performance. FAH-QLoRA integrates two key techniques: i) dynamically adjusting LoRA ranks across training rounds to promote faster convergence and lower resource usage, and ii) assigning heterogeneous LoRA ranks across devices to mitigate the straggler effect during the FL process, ensuring that heterogeneous resource constraints are met. We analyze the convergence of FAH-QLoRA under general non-convex and non-IID FL settings. Extensive experiments demonstrate that FAH-QLoRA can reduce training time by up to 45.86% and memory usage by up to 44.15% compared to baseline methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 30adb4ae-883a-4515-806d-42da166feb3dCited by top-tier papers5
- PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud CollaborationYi Liu, Weixiang Han, Chengjun Cai, Xingliang Yuan et al.INFOCOM 2026 · 2 citations
- FedKRSO: Communication and Memory Efficient Federated Fine-Tuning of Large Language ModelsGuohao Yang, Tongle Wu, Yuanxiong Guo, Ying Sun et al.INFOCOM 2026 · 2 citations
- WinFLoRA: Incentivizing Client-Adaptive Aggregation in Federated LoRA under Privacy HeterogeneityMengsha Kou, Xiaoyu Xia, Ziqi Wang, Ibrahim Khalil et al.WWW 2026 · 1 citation
- Dual-Phase Federated Deep Unlearning via Weight-Aware Rollback and ReconstructionChangjun Zhou, Jintao Zheng, Leyou Yang, Pengfei WangINFOCOM 2026
- Heterogeneity-Oblivious Robust Federated LearningWeiyao Zhang, Jinyang Li, Qi Song, Miao Wang et al.INFOCOM 2026
Related papers
- Heterogeneous Federated Fine-Tuning with Parallel One-Rank AdaptationZikai Zhang, Rui Hu, Jiahao XuICLR 2026 · 6 citations
- Towards Robust Parameter-Efficient Fine-Tuning for Federated LearningXiuwen Fang, Mang YeNeurIPS 2025 · 2 citations
- Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert AssemblyYebo Wu, Jingguang Li, Zhijiang Guo, Li LiICML 2026
- Tensor-Aggregated LoRA in Federated Fine-TuningZhixuan Li, Binqian Xu, Xiangbo Shu, Jiachao Zhang et al.ICCV 2025 · 2 citations
- Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation ModelsYae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi et al.EMNLP 2024 · 36 citations
