Local Superior Soups: A Catalyst for Model Merging in Cross-Silo Federated Learning
Minghui Chen, Meirui Jiang, Xin Zhang, Qi Dou, Zehua Wang, Xiaoxiao Li
摘要
Federated learning (FL) is a learning paradigm that enables collaborative training of models using decentralized data. Recently, the utilization of pre-trained weight initialization in FL has been demonstrated to effectively improve model performance. However, the evolving complexity of current pre-trained models, characterized by a substantial increase in parameters, markedly intensifies the challenges associated with communication rounds required for their adaptation to FL. To address these communication cost issues and increase the performance of pre-trained model adaptation in FL, we propose an innovative model interpolation-based local training technique called ``Local Superior Soups.'' Our method enhances local training across different clients, encouraging the exploration of a connected low-loss basin within a few communication rounds through regularized model interpolation. This approach acts as a catalyst for the seamless adaptation of pre-trained models in in FL. We demonstrated its effectiveness and efficiency across diverse widely-used FL datasets. Our code is available at https://github.com/ubc-tea/Local-Superior-Soups.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FedMerge: Federated Model Merging for PersonalizationShutong Chen, Tianyi Zhou, Guodong Long, Jing Jiang 等AAAI 2026 · 被引用 2 次
- Self-Soupervision: Cooking Model Soups without LabelsAnthony Fuller, James Green, Evan ShelhamerICML 2026
- Can Textual Gradient Work in Federated Learning?Minghui Chen, Ruinan Jin, Wenlong Deng, Yuanyuan Chen 等ICLR 2025
- Modeling Multi-Task Model Merging as Adaptive Projective Gradient DescentYongxian Wei, Anke Tang, Li Shen, Zixuan Hu 等ICML 2025
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
相关 Paper
- On the Importance and Applicability of Pre-Training for Federated LearningHong-You Chen, Cheng-Hao Tu, Ziwei Li, Han-Wei Shen 等ICLR 2023 · 被引用 10 次
- Rethinking the Starting Point: Collaborative Pre-Training for Federated Downstream TasksYun-Wei Chu, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. BrintonAAAI 2025 · 被引用 1 次
- Guiding The Last Layer in Federated Learning with Pre-Trained ModelsGwen Legate, Nicolas Bernier, Lucas Page-Caccia, Edouard Oyallon 等NeurIPS 2023 · 被引用 31 次
- Graph Ladling: Shockingly Simple Parallel GNN Training without Intermediate CommunicationAjay Kumar Jaiswal, Shiwei Liu, Tianlong Chen, Ying Ding 等ICML 2023 · 被引用 8 次
- Towards Instance-adaptive Inference for Federated LearningChun-Mei Feng, Kai Yu, Nian Liu, Xinxing Xu 等ICCV 2023 · 被引用 18 次
