pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
Sajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin Pedarsani
摘要
Vision-Language Models (VLMs) like CLIP have demonstrated remarkable generalization in zero-and few-shot settings, but adapting them efficiently to decentralized, heterogeneous data remains a challenge. While prompt tuning has emerged as a popular parameter-efficient approach in personalized federated learning, existing methods often sacrifice generalization in favor of personalization, struggling particularly on unseen classes or domains. In this work, we propose pFedMMA, a personalized federated learning framework that leverages multi-modal adapters for vision-language tasks. Each adapter contains modality-specific up-and downprojection layers alongside a globally shared projection that aligns cross-modal features. Our optimization strategy allows clients to locally adapt to personalized data distributions while collaboratively training the shared projection to improve global generalization. This design is also communication-efficient, as only the shared component is exchanged during communication rounds. Through extensive experiments across eleven datasets, including domain-and label-shift scenarios, we show that pFedMMA achieves state-of-the-art trade-offs between personalization and generalization, outperforming recent federated prompt tuning methods. Code is available at https://github.com/sajjad-ucsb/pFedMMA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Training-Free Adversarial Robustness in Computational MRIMahdi Saberi, Chi Zhang, Mehmet AkcakayaICML 2026 · 被引用 2 次
- Beyond Description: Federated Adaptation via Semantic-Visual Prototype AlignmentJiarong Yang, Yuan LiuICML 2026
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
相关 Paper
- Harmonizing Generalization and Personalization in Federated Prompt LearningTianyu Cui, Hongxia Li, Jingya Wang, Ye ShiICML 2024 · 被引用 31 次
- FedPHA: Federated Prompt Learning for Heterogeneous Client AdaptationChengying Fang, Wenke Huang, Guancheng Wan, Yihao Yang 等ICML 2025
- FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated LearningYubin Zheng, Pak-Hei Yeung, Jing Xia, Tianjie Ju 等ACM MM 2025
- pFedPrompt: Learning Personalized Prompt for Vision-Language Models in Federated LearningTao Guo, Song Guo, Junxiao WangWWW 2023 · 被引用 101 次
- Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language ModelsJun Luo, Chen Chen, Shandong WuICLR 2025
