Topology-aware Generalization of Decentralized SGD
Tongtian Zhu, Fengxiang He, Lan Zhang, Zhengyang Niu, Mingli Song, Dacheng Tao
Abstract
This paper studies the algorithmic stability and generalizability of decentralized stochastic gradient descent (D-SGD). We prove that the consensus model learned by D-SGD is -stable in expectation in the non-convex non-smooth setting, where is the total sample size, is the worker number, and is the spectral gap that measures the connectivity of the communication topology. These results then deliver an in-average generalization bound, which is non-vacuous even when is closed to , in contrast to vacuous as suggested by existing literature on the projected version of D-SGD. Our theory indicates that the generalizability of D-SGD is positively correlated with the spectral gap, and can explain why consensus control in initial training phase can ensure better generalization. Experiments of VGG-11 and ResNet-18 on CIFAR-10, CIFAR-100 and Tiny-ImageNet justify our theory. To our best knowledge, this is the first work on the topology-aware generalization of vanilla D-SGD. Code is available at https://github.com/Raiden-Zhu/Generalization-of-DSGD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse TrainingRong Dai, Li Shen, Fengxiang He, Xinmei Tian et al.ICML 2022 · 163 citations
- Improving the Model Consistency of Decentralized Federated LearningYifan Shi, Li Shen, Kang Wei, Yan Sun et al.ICML 2023 · 89 citations
- Stability-Based Generalization Analysis of the Asynchronous Decentralized SGDXiaoge Deng, Tao Sun, Shengwei Li, Dongsheng LiAAAI 2023 · 26 citations
- Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time ConvergenceYuki Takezawa, Ryoma Sato, Han Bao, Kenta Niwa et al.NeurIPS 2023 · 22 citations
- Decentralized SGD and Average-direction SAM are Asymptotically EquivalentTongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song et al.ICML 2023 · 21 citations
Builds on18
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 1,822 citations
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi et al.ICML 2020 · 623 citations
- Stability of Stochastic Gradient Descent on Nonsmooth Convex LossesRaef Bassily, Vitaly Feldman, Cristóbal Guzmán, Kunal TalwarNeurIPS 2020 · 240 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
- Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient DescentYunwen Lei, Yiming YingICML 2020 · 165 citations
Related papers
- Improved Stability and Generalization Guarantees of the Decentralized SGD AlgorithmBatiste Le Bars, Aurélien Bellet, Marc Tommasi, Kevin Scaman et al.ICML 2024 · 13 citations
- Stability and Generalization of the Decentralized Stochastic Gradient Descent Ascent AlgorithmMiaoxi Zhu, Li Shen, Bo Du, Dacheng TaoNeurIPS 2023 · 12 citations
- Improved Convergence Analysis of Topology Dependence in Decentralized SGDYuki Takezawa, Anastasiia Koloskova, Sebastian StichICML 2026
- Stability and Generalization of Decentralized Stochastic Gradient DescentTao Sun, Dongsheng Li, Bao WangAAAI 2021 · 44 citations
- Beyond spectral gap: the role of the topology in decentralized learningThijs Vogels, Hadrien Hendrikx, Martin JaggiNeurIPS 2022 · 48 citations
