pFedPrompt: Learning Personalized Prompt for Vision-Language Models in Federated Learning
Tao Guo, Song Guo, Junxiao Wang
Abstract
Pre-trained vision-language models like CLIP show great potential in learning representations that capture latent characteristics of users. A recently proposed method called Contextual Optimization (CoOp) introduces the concept of training prompt for adapting pre-trained vision-language models. Given the lightweight nature of this method, researchers have migrated the paradigm from centralized to decentralized system to innovate the collaborative training framework of Federated Learning (FL). However, current prompt training in FL mainly focuses on modeling user consensus and lacks the adaptation to user characteristics, leaving the personalization of prompt largely under-explored. Researches over the past few years have applied personalized FL (pFL) approaches to customizing models for heterogeneous users. Unfortunately, we find that with the variation of modality and training behavior, directly applying the pFL methods to prompt training leads to insufficient personalization and performance. To bridge the gap, we present pFedPrompt, which leverages the unique advantage of multimodality in vision-language models by learning user consensus from linguistic space and adapting to user characteristics in visual space in a non-parametric manner. Through this dual collaboration, the learned prompt will be fully personalized and aligned to the user’s local characteristics. We conduct extensive experiments across various datasets under the FL setting with statistical heterogeneity. The results demonstrate the superiority of our pFedPrompt against the alternative approaches with robust performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eabc09a2-6386-44f2-a142-7b06c2a93033Cited by top-tier papers32
- FedDAT: An Approach for Foundation Model Finetuning in Multi-Modal Heterogeneous Federated LearningHaokun Chen, Yao Zhang, Denis Krompass, Jindong Gu et al.AAAI 2024 · 105 citations
- Faithful Vision-Language Interpretation via Concept Bottleneck ModelsSongning Lai, Lijie Hu, Junxiao Wang, Laure Berti-Équille et al.ICLR 2024 · 42 citations
- Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation ModelsYae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi et al.EMNLP 2024 · 36 citations
- Federated Text-driven Prompt Generation for Vision-Language ModelsChen Qiu, Xingyu Li, Chaithanya Kumar Mummadi, Madan Ravi Ganesh et al.ICLR 2024 · 33 citations
- Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and MethodBikang Pan, Wei Huang, Ye ShiNeurIPS 2024 · 28 citations
Related papers
- Harmonizing Generalization and Personalization in Federated Prompt LearningTianyu Cui, Hongxia Li, Jingya Wang, Ye ShiICML 2024 · 31 citations
- FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language ModelsXinting Liao, Weiming Liu, Jiaming Qian, Pengyang Zhou et al.ICML 2025
- Global and Local Prompts Cooperation via Optimal Transport for Federated LearningHongxia Li, Wei Huang, Jingya Wang, Ye ShiCVPR 2024
- Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language ModelsJun Luo, Chen Chen, Shandong WuICLR 2025
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
