Consensus Control for Decentralized Deep Learning
Lingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi, Sebastian U. Stich
摘要
Decentralized training of deep learning models enables on-device learning over networks, as well as efficient scaling to large compute clusters. Experiments in earlier works reveal that, even in a data-center setup, decentralized training often suffers from the degradation in the quality of the model: the training and test performance of models trained in a decentralized fashion is in general worse than that of models trained in a centralized fashion, and this performance drop is impacted by parameters such as network size, communication topology and data partitioning. We identify the changing consensus distance between devices as a key parameter to explain the gap between centralized and decentralized training. We show in theory that when the training consensus distance is lower than a critical quantity, decentralized training converges as fast as the centralized counterpart. We empirically validate that the relation between generalization performance and consensus distance is consistent with this theoretical observation. Our empirical insights allow the principled design of better decentralized training schemes that mitigate the performance drop. To this end, we provide practical training guidelines and exemplify its effectiveness on the data-center setup as the important first step.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- On Bridging Generic and Personalized Federated Learning for Image ClassificationHong-You Chen, Wei-Lun ChaoICLR 2022 · 被引用 329 次
- An Improved Analysis of Gradient Tracking for Decentralized Machine LearningAnastasia Koloskova, Tao Lin, Sebastian U. StichNeurIPS 2021 · 被引用 148 次
- Exponential Graph is Provably Efficient for Decentralized Deep TrainingBicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu 等NeurIPS 2021 · 被引用 123 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang 等ICCV 2021 · 被引用 73 次
它引用的顶会 Paper10
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 被引用 672 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 被引用 263 次
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 被引用 199 次
相关 Paper
- From Promise to Practice: Realizing High-performance Decentralized TrainingZesen Wang, Jiaojiao Zhang, Xuyang Wu, Mikael JohanssonICLR 2025
- Beyond spectral gap: the role of the topology in decentralized learningThijs Vogels, Hadrien Hendrikx, Martin JaggiNeurIPS 2022 · 被引用 48 次
- Stability and Generalization of Decentralized Stochastic Gradient DescentTao Sun, Dongsheng Li, Bao WangAAAI 2021 · 被引用 44 次
- Topology-aware Generalization of Decentralized SGDTongtian Zhu, Fengxiang He, Lan Zhang, Zhengyang Niu 等ICML 2022 · 被引用 58 次
- Unveiling the Power of Multiple Gossip Steps: A Stability-Based Generalization Analysis in Decentralized TrainingQinglun Li, Yingqi Liu, Miao Zhang, Xiaochun Cao 等NeurIPS 2025 · 被引用 3 次
