Enhancing Storage and Computational Efficiency in Federated Multimodal Learning for Large-Scale Models
Zixin Zhang, Fan Qi, Changsheng Xu
摘要
The remarkable generalization of large-scale models has recently gained significant attention in multimodal research. However, deploying heterogeneous large-scale models with different modalities under Federated Learning (FL) to protect data privacy imposes tremendous challenges on clients' limited computation and storage. In this work, we propose M 2 FEDSA to address the above issue. We realize modularized decomposition of large-scale models via Split Learning (SL) and only retain privacy-sensitive modules on clients, alleviating storage overhead. By freezing largescale models and introducing two specialized lightweight adapters, the models can better focus on task-specific knowledge and enhance modalityspecific knowledge, improving the model's adaptability to different tasks while balancing efficiency. In addition, M 2 FEDSA further improves performance by transferring multimodal knowledge to unimodal clients at both the feature and decision levels, which leverages the complementarity of different modalities. Extensive experiments on various multimodal classification tasks validate the effectiveness of our proposed M 2 FEDSA. The code is made available publicly at https: //github.com/M2FedSA/M-2FedSA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- FedAFD: Multimodal Federated Learning via Adversarial Fusion and DistillationMin Tan, Junchao Ma, Yinfu FENG, Jiajun Ding 等CVPR 2026 · 被引用 1 次
- MFC: Mixed Federated Clustering based on Cross-modal Feature DecouplingXiaxia He, Boyue Wang, Junbin Gao, Yongli Hu 等KDD 2026
- FedMSplit: Correlation-Adaptive Federated Multi-Task Learning across Multimodal Split NetworksJiayi Chen, Aidong ZhangKDD 2022 · 被引用 86 次
- FedDAT: An Approach for Foundation Model Finetuning in Multi-Modal Heterogeneous Federated LearningHaokun Chen, Yao Zhang, Denis Krompass, Jindong Gu 等AAAI 2024 · 被引用 105 次
- UniFLoW: Universal Multi-Modal Federated LoRA Fine-Tuning Framework with Analytical Aggregationhaoyuan liang, Zhiyu Ye, Jielong Tang, Yang Yang 等ICML 2026
