Local Superior Soups: A Catalyst for Model Merging in Cross-Silo Federated Learning
Minghui Chen, Meirui Jiang, Xin Zhang, Qi Dou, Zehua Wang, Xiaoxiao Li
Abstract
Federated learning (FL) is a learning paradigm that enables collaborative training of models using decentralized data. Recently, the utilization of pre-trained weight initialization in FL has been demonstrated to effectively improve model performance. However, the evolving complexity of current pre-trained models, characterized by a substantial increase in parameters, markedly intensifies the challenges associated with communication rounds required for their adaptation to FL. To address these communication cost issues and increase the performance of pre-trained model adaptation in FL, we propose an innovative model interpolation-based local training technique called ``Local Superior Soups.'' Our method enhances local training across different clients, encouraging the exploration of a connected low-loss basin within a few communication rounds through regularized model interpolation. This approach acts as a catalyst for the seamless adaptation of pre-trained models in in FL. We demonstrated its effectiveness and efficiency across diverse widely-used FL datasets. Our code is available at https://github.com/ubc-tea/Local-Superior-Soups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59a88e68-1fbf-4844-9ad9-64b74fd72fa2Cited by top-tier papers4
- FedMerge: Federated Model Merging for PersonalizationShutong Chen, Tianyi Zhou, Guodong Long, Jing Jiang et al.AAAI 2026 · 2 citations
- Self-Soupervision: Cooking Model Soups without LabelsAnthony Fuller, James Green, Evan ShelhamerICML 2026
- Can Textual Gradient Work in Federated Learning?Minghui Chen, Ruinan Jin, Wenlong Deng, Yuanyuan Chen et al.ICLR 2025
- Modeling Multi-Task Model Merging as Adaptive Projective Gradient DescentYongxian Wei, Anke Tang, Li Shen, Zixuan Hu et al.ICML 2025
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
Related papers
- On the Importance and Applicability of Pre-Training for Federated LearningHong-You Chen, Cheng-Hao Tu, Ziwei Li, Han-Wei Shen et al.ICLR 2023 · 10 citations
- Rethinking the Starting Point: Collaborative Pre-Training for Federated Downstream TasksYun-Wei Chu, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. BrintonAAAI 2025 · 1 citation
- Guiding The Last Layer in Federated Learning with Pre-Trained ModelsGwen Legate, Nicolas Bernier, Lucas Page-Caccia, Edouard Oyallon et al.NeurIPS 2023 · 31 citations
- Graph Ladling: Shockingly Simple Parallel GNN Training without Intermediate CommunicationAjay Kumar Jaiswal, Shiwei Liu, Tianlong Chen, Ying Ding et al.ICML 2023 · 8 citations
- Towards Instance-adaptive Inference for Federated LearningChun-Mei Feng, Kai Yu, Nian Liu, Xinxing Xu et al.ICCV 2023 · 18 citations
