Topology-aware Generalization of Decentralized SGD
Tongtian Zhu, Fengxiang He, Lan Zhang, Zhengyang Niu, Mingli Song, Dacheng Tao
摘要
This paper studies the algorithmic stability and generalizability of decentralized stochastic gradient descent (D-SGD). We prove that the consensus model learned by D-SGD is -stable in expectation in the non-convex non-smooth setting, where is the total sample size, is the worker number, and is the spectral gap that measures the connectivity of the communication topology. These results then deliver an in-average generalization bound, which is non-vacuous even when is closed to , in contrast to vacuous as suggested by existing literature on the projected version of D-SGD. Our theory indicates that the generalizability of D-SGD is positively correlated with the spectral gap, and can explain why consensus control in initial training phase can ensure better generalization. Experiments of VGG-11 and ResNet-18 on CIFAR-10, CIFAR-100 and Tiny-ImageNet justify our theory. To our best knowledge, this is the first work on the topology-aware generalization of vanilla D-SGD. Code is available at https://github.com/Raiden-Zhu/Generalization-of-DSGD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse TrainingRong Dai, Li Shen, Fengxiang He, Xinmei Tian 等ICML 2022 · 被引用 163 次
- Improving the Model Consistency of Decentralized Federated LearningYifan Shi, Li Shen, Kang Wei, Yan Sun 等ICML 2023 · 被引用 89 次
- Stability-Based Generalization Analysis of the Asynchronous Decentralized SGDXiaoge Deng, Tao Sun, Shengwei Li, Dongsheng LiAAAI 2023 · 被引用 26 次
- Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time ConvergenceYuki Takezawa, Ryoma Sato, Han Bao, Kenta Niwa 等NeurIPS 2023 · 被引用 22 次
- Decentralized SGD and Average-direction SAM are Asymptotically EquivalentTongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song 等ICML 2023 · 被引用 21 次
它引用的顶会 Paper18
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 被引用 1,822 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Stability of Stochastic Gradient Descent on Nonsmooth Convex LossesRaef Bassily, Vitaly Feldman, Cristóbal Guzmán, Kunal TalwarNeurIPS 2020 · 被引用 240 次
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 被引用 199 次
- Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient DescentYunwen Lei, Yiming YingICML 2020 · 被引用 165 次
相关 Paper
- Improved Stability and Generalization Guarantees of the Decentralized SGD AlgorithmBatiste Le Bars, Aurélien Bellet, Marc Tommasi, Kevin Scaman 等ICML 2024 · 被引用 13 次
- Stability and Generalization of the Decentralized Stochastic Gradient Descent Ascent AlgorithmMiaoxi Zhu, Li Shen, Bo Du, Dacheng TaoNeurIPS 2023 · 被引用 12 次
- Improved Convergence Analysis of Topology Dependence in Decentralized SGDYuki Takezawa, Anastasiia Koloskova, Sebastian StichICML 2026
- Stability and Generalization of Decentralized Stochastic Gradient DescentTao Sun, Dongsheng Li, Bao WangAAAI 2021 · 被引用 44 次
- Beyond spectral gap: the role of the topology in decentralized learningThijs Vogels, Hadrien Hendrikx, Martin JaggiNeurIPS 2022 · 被引用 48 次
