Enhancing Storage and Computational Efficiency in Federated Multimodal Learning for Large-Scale Models
Zixin Zhang, Fan Qi, Changsheng Xu
Abstract
The remarkable generalization of large-scale models has recently gained significant attention in multimodal research. However, deploying heterogeneous large-scale models with different modalities under Federated Learning (FL) to protect data privacy imposes tremendous challenges on clients' limited computation and storage. In this work, we propose M 2 FEDSA to address the above issue. We realize modularized decomposition of large-scale models via Split Learning (SL) and only retain privacy-sensitive modules on clients, alleviating storage overhead. By freezing largescale models and introducing two specialized lightweight adapters, the models can better focus on task-specific knowledge and enhance modalityspecific knowledge, improving the model's adaptability to different tasks while balancing efficiency. In addition, M 2 FEDSA further improves performance by transferring multimodal knowledge to unimodal clients at both the feature and decision levels, which leverages the complementarity of different modalities. Extensive experiments on various multimodal classification tasks validate the effectiveness of our proposed M 2 FEDSA. The code is made available publicly at https: //github.com/M2FedSA/M-2FedSA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9c00cc9-01f0-4f3c-b3a9-da3e410ec04eCited by top-tier papers1
Ask how each one uses itBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- FedAFD: Multimodal Federated Learning via Adversarial Fusion and DistillationMin Tan, Junchao Ma, Yinfu FENG, Jiajun Ding et al.CVPR 2026 · 1 citation
- MFC: Mixed Federated Clustering based on Cross-modal Feature DecouplingXiaxia He, Boyue Wang, Junbin Gao, Yongli Hu et al.KDD 2026
- FedMSplit: Correlation-Adaptive Federated Multi-Task Learning across Multimodal Split NetworksJiayi Chen, Aidong ZhangKDD 2022 · 86 citations
- FedDAT: An Approach for Foundation Model Finetuning in Multi-Modal Heterogeneous Federated LearningHaokun Chen, Yao Zhang, Denis Krompass, Jindong Gu et al.AAAI 2024 · 105 citations
- UniFLoW: Universal Multi-Modal Federated LoRA Fine-Tuning Framework with Analytical Aggregationhaoyuan liang, Zhiyu Ye, Jielong Tang, Yang Yang et al.ICML 2026
