DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning Models
Nastaran Saadati, Minh Pham, Nasla Saleem, Joshua R. Waite, Aditya Balu, Zhanhong Jiang, Chinmay Hegde, Soumik Sarkar
Abstract
Recent advances in decentralized deep learning algorithms have demonstrated cutting-edge performance on various tasks with large pretrained models. However, a pivotal prerequisite for achieving this level of competitiveness is the significant communication and computation overheads when updating these models, which prohibits the applications of them to real-world scenarios. To address this issue, drawing inspiration from advanced model merging techniques without requiring additional training, we introduce the Decentralized Iterative Merging-And-Training (DIMAT) paradigm-a novel decentralized deep learning framework. Within DIMAT, each agent is trained on their local data and periodically merged with their neighboring agents using advanced model merging techniques like activation matching until convergence is achieved. DIMAT provably converges with the best available rate for non-convex functions with various first-order methods, while yielding tighter error bounds compared to the popular existing approaches. We conduct a comprehensive empirical analysis to validate DIMAT's superiority over baselines across diverse computer vision tasks sourced from multiple datasets. Empirical results validate our theoretical claims by showing that DIMAT attains faster and higher initial gain in accuracy with independent and identically distributed (IID) and non-IID data, incurring lower communication overhead. This DIMAT paradigm presents a new op-portunity for the future decentralized learning, enhancing its adaptability to real-world with sparse and lightweight communication and computation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffc8ee67-6fd0-4b07-a480-88229959798cBuilds on20
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- ZipIt! Merging Models from Different Tasks without TrainingGeorge Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh et al.ICLR 2024 · 185 citations
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 118 citations
Related papers
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang et al.ICCV 2021 · 73 citations
- Cross-Gradient Aggregation for Decentralized Learning from Non-IID DataYasaman Esfandiari, Sin Yong Tan, Zhanhong Jiang, Aditya Balu et al.ICML 2021 · 61 citations
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 95 citations
- A Decentralized Parallel Algorithm for Training Generative Adversarial NetsMingrui Liu, Wei Zhang, Youssef Mroueh, Xiaodong Cui et al.NeurIPS 2020 · 6 citations
- On The Surprising Effectiveness of a Single Global Merging in Decentralized LearningTongtian Zhu, Tianyu Zhang, Mingze Wang, Zhanpeng Zhou et al.ICLR 2026 · 2 citations
