Overcoming Heterogeneous Data in Federated Medical Vision-Language Pre-training: A Triple-Embedding Model Selector Approach
Aowen Wang, Zhiwang Zhang, Dongang Wang, Fanyi Wang, Haotian Hu, Jinyang Guo, Yipeng Zhou, Chaoyi Pang, Shiting Wen
Abstract
The scarcity of data in the medical field brings challenges to collaborative training in medical vision-language pre-training (VLP) across different clients Thus, collaborative training in medical VLP faces two significant challenges: First, the medical data requires privacy and therefore cannot be directly shared across different clients. Second, medical data distribution across institutes is typically heterogeneous, hindering local model alignment and representation capabilities. To simultaneously overcome these two challenges, we propose a framework called personalized model selector with fused multimodal information (PMS-FM). The contribution of PMS-FM is two-fold: 1) PMS-FM uses embeddings to represent information in different formats, allowing for the fusion of multimodal data. 2) PMS-FM adapts to personalized data distributions by training multiple models. A model selector then identifies and selects the best-performing model for each individual client. Extensive experiments with multiple real-world medical datasets demonstrate the superb performance of PMS-FM over existing federated learning methods on different zero-shot classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da44eba9-6891-4630-a0ee-44da095b98d0Cited by top-tier papers1
Ask how each one uses itBuilds on8
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- FedALA: Adaptive Local Aggregation for Personalized Federated LearningJianqing Zhang, Yang Hua, Hao Wang, Tao Song et al.AAAI 2023 · 445 citations
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti et al.NeurIPS 2022 · 302 citations
- FedFed: Feature Distillation against Data Heterogeneity in Federated LearningZhiqin Yang, Yonggang Zhang, Yu Zheng, Xinmei Tian et al.NeurIPS 2023 · 166 citations
- Personalized Federated Learning through Local MemorizationOthmane Marfoq, Giovanni Neglia, Richard Vidal, Laetitia KameniICML 2022 · 124 citations
Related papers
- Federated CLIP for Resource-Efficient Heterogeneous Medical Image ClassificationYihang Wu, Ahmad ChaddadAAAI 2026 · 1 citation
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language ModelsSajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin PedarsaniICLR 2026 · 3 citations
- Beyond Description: Federated Adaptation via Semantic-Visual Prototype AlignmentJiarong Yang, Yuan LiuICML 2026
- MH-pFLID: Model Heterogeneous personalized Federated Learning via Injection and Distillation for Medical Data AnalysisLuyuan Xie, Manqing Lin, Tianyu Luan, Cong Li et al.ICML 2024 · 21 citations
- Decoupled Training with Local Reinforcement Fine-Tuning in Federated LearningYuting Ma, Lechao Cheng, Xiaohua XuICML 2026
