Beyond Description: Federated Adaptation via Semantic-Visual Prototype Alignment
Jiarong Yang, Yuan Liu
Abstract
Adopting pre-trained Vision-Language Models (VLMs) in Federated Learning (FL) presents a promising avenue for mitigating data scarcity and heterogeneity. However, existing solutions suffer from high computational complexity or ineffective knowledge aggregation. To address these problems, we propose FedSPA (Federated Adaptation via Semantic-Visual Prototype Alignment). On the client side, FedSPA restricts local optimization to visual prototypes, enabling lightweight personalization. On the server side, we introduce a semantic alignment module that leverages client-uploaded prototypes to minimize a contrastive objective, aligning global semantic prototypes with heterogeneous visual distributions and thereby shifting the paradigm from traditional "learning-to-describe" (optimizing static prompts) to "learning-to-align". Extensive experiments demonstrate that FedSPA significantly outperforms state-of-the-art methods in both personalized and global benchmarks, while substantially reducing computational overhead. The code is available at https://github.com/ eejiarong/FedSPA-main.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d28b0fb3-91da-4e8a-b0ab-caa1764a03ebBuilds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Personalized Federated Learning with Moreau EnvelopesCanh T. Dinh, Nguyen Hoang Tran, Tuan Dung NguyenNeurIPS 2020 · 1,542 citations
Related papers
- FedPHA: Federated Prompt Learning for Heterogeneous Client AdaptationChengying Fang, Wenke Huang, Guancheng Wan, Yihao Yang et al.ICML 2025
- Enhancing Visual Representation with Textual Semantics: Textual Semantics-Powered Prototypes for Heterogeneous Federated LearningXinghao Wu, Jianwei Niu, Xuefeng Liu, Guogang Zhu et al.CVPR 2026 · 4 citations
- Harmonizing Generalization and Personalization in Federated Prompt LearningTianyu Cui, Hongxia Li, Jingya Wang, Ye ShiICML 2024 · 31 citations
- TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language ModelsLi Zhang, Zhongxuan Han, Xiaohua Feng, Jiaming Zhang et al.AAAI 2026 · 1 citation
- Decoupled Training with Local Reinforcement Fine-Tuning in Federated LearningYuting Ma, Lechao Cheng, Xiaohua XuICML 2026
