DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning Models
Nastaran Saadati, Minh Pham, Nasla Saleem, Joshua R. Waite, Aditya Balu, Zhanhong Jiang, Chinmay Hegde, Soumik Sarkar
摘要
Recent advances in decentralized deep learning algorithms have demonstrated cutting-edge performance on various tasks with large pretrained models. However, a pivotal prerequisite for achieving this level of competitiveness is the significant communication and computation overheads when updating these models, which prohibits the applications of them to real-world scenarios. To address this issue, drawing inspiration from advanced model merging techniques without requiring additional training, we introduce the Decentralized Iterative Merging-And-Training (DIMAT) paradigm-a novel decentralized deep learning framework. Within DIMAT, each agent is trained on their local data and periodically merged with their neighboring agents using advanced model merging techniques like activation matching until convergence is achieved. DIMAT provably converges with the best available rate for non-convex functions with various first-order methods, while yielding tighter error bounds compared to the popular existing approaches. We conduct a comprehensive empirical analysis to validate DIMAT's superiority over baselines across diverse computer vision tasks sourced from multiple datasets. Empirical results validate our theoretical claims by showing that DIMAT attains faster and higher initial gain in accuracy with independent and identically distributed (IID) and non-IID data, incurring lower communication overhead. This DIMAT paradigm presents a new op-portunity for the future decentralized learning, enhancing its adaptability to real-world with sparse and lightweight communication and computation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li 等CVPR 2022 · 被引用 364 次
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 被引用 301 次
- ZipIt! Merging Models from Different Tasks without TrainingGeorge Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh 等ICLR 2024 · 被引用 185 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
相关 Paper
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang 等ICCV 2021 · 被引用 73 次
- Cross-Gradient Aggregation for Decentralized Learning from Non-IID DataYasaman Esfandiari, Sin Yong Tan, Zhanhong Jiang, Aditya Balu 等ICML 2021 · 被引用 61 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
- A Decentralized Parallel Algorithm for Training Generative Adversarial NetsMingrui Liu, Wei Zhang, Youssef Mroueh, Xiaodong Cui 等NeurIPS 2020 · 被引用 6 次
- On The Surprising Effectiveness of a Single Global Merging in Decentralized LearningTongtian Zhu, Tianyu Zhang, Mingze Wang, Zhanpeng Zhou 等ICLR 2026 · 被引用 2 次
