Decentralized Stochastic Nonconvex Optimization under the (L0, L1)-Smoothness
Luo Luo, Xue Cui, Tingkai Jia, Cheng Chen
摘要
This paper focuses on the decentralized stochastic optimization problem f(x)=1/m ∑i=1m fi(x) over a connected network of n agents, where each local function has the form of fi(x) = E[F(x;ξ i)] which satisfies the (L0,L1)-smooth condition but possibly nonconvex and each random variable ξi follows distribution Di. We propose a novel algorithm called decentralized normalized stochastic gradient descent (DNSGD), which can achieve an ε-stationary point at each local agent. We present a new framework for analyzing decentralized first-order methods in the (L0,L1)-smooth setting, based on the Lyapunov function related to the product of the gradient norm and the consensus error. We show that the proposed algorithm attains the upper bounds on the sample complexity of O(m-1 (Lfσ2Δfε-4 + σ2ε-2 + Lf-2 L13σ2?fΔ-1 + Lf-2 L12σ2)) per agent and the communication complexity of O ((Lfε-2 + L1ε-1)γ-1/2 Δf), where Lf=L0 +L1ζ, σ2 is the variance of the stochastic gradient, Δf is the initial optimal function value gap, γ is the spectral gap of the network, and ζ is the degree of the gradient dissimilarity. In the special case of L1=0, the above results (nearly) match the lower bounds of decentralized stochastic nonconvex optimization under the standard smoothness. We also conduct numerical experiments to show the empirical superiority of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 被引用 231 次
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 被引用 177 次
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 被引用 139 次
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 被引用 132 次
相关 Paper
- Decentralized Riemannian Gradient Descent on the Stiefel ManifoldShixiang Chen, Alfredo García, Mingyi Hong, Shahin ShahrampourICML 2021 · 被引用 64 次
- On the Complexity of Finite-Sum Smooth Optimization under the Polyak-Łojasiewicz ConditionYunyan Bai, Yuxing Liu, Luo LuoICML 2024 · 被引用 2 次
- Topology-aware Generalization of Decentralized SGDTongtian Zhu, Fengxiang He, Lan Zhang, Zhengyang Niu 等ICML 2022 · 被引用 58 次
- Efficient Decentralized Stochastic Gradient Descent Method for Nonconvex Finite-Sum Optimization ProblemsWenkang Zhan, Gang Wu, Hongchang GaoAAAI 2022 · 被引用 8 次
- A Hybrid Variance-Reduced Method for Decentralized Stochastic Non-Convex OptimizationRan Xin, Usman A. Khan, Soummya KarICML 2021 · 被引用 51 次
