Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning
Yuting Ma, Lechao Cheng, Xiaohua Xu
Abstract
Federated Learning (FL) with pre-trained Vision-Language Models (VLMs) has emerged as a promising paradigm for various downstream tasks. By leveraging its strong representations, recent studies improve task adaptation under insufficient local data while preserving generalization. However, these methods emphasize fully local optimization with simple parameter aggregation, which can amplify inter-client optimization inconsistency and intra-client over-specialization under heterogeneous and full-data FL settings, making it difficult to balance global task adaptation and generalization. To address these challenges, we propose FedDTL, a novel federated VLM framework that decouples the image encoder and text encoder across clients and the server. Through decoupled encoder training with server-client modality alignment, FedDTL promotes coherent global semantic update and reduces inter-client optimization inconsistency, improving global task adaptation. To further mitigate intra-client over-specialization, we introduce a two-stage local fine-tuning, where a supervised fine-tuning stage enables rapid and reliable warm-start, followed by a reinforcement learning stage that enhances generalization. Extensive experiments on multiple benchmarks, including label skew and feature shift, demonstrate that FedDTL achieves an effective balance between global task adaptation and generalization under various FL data distributions in both few-shot and full-data regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c8d54eb-be67-463d-b40f-78dec5e1b482Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- FedCDA: Federated Learning with Cross-rounds Divergence-aware AggregationHaozhao Wang, Haoran Xu, Yichen Li, Yuan Xu et al.ICLR 2024 · 62 citations
Related papers
- TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language ModelsLi Zhang, Zhongxuan Han, Xiaohua Feng, Jiaming Zhang et al.AAAI 2026 · 1 citation
- Beyond Description: Federated Adaptation via Semantic-Visual Prototype AlignmentJiarong Yang, Yuan LiuICML 2026
- Federated Disentangled Tuning with Textual Prior Decoupling and Visual Dynamic AdaptationYihao Yang, Wenke Huang, Guancheng Wan, Bin Yang et al.ICML 2025
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language ModelsSajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin PedarsaniICLR 2026 · 3 citations
- FedPHA: Federated Prompt Learning for Heterogeneous Client AdaptationChengying Fang, Wenke Huang, Guancheng Wan, Yihao Yang et al.ICML 2025
