Decentralized Stochastic Nonconvex Optimization under the (L0, L1)-Smoothness
Luo Luo, Xue Cui, Tingkai Jia, Cheng Chen
Abstract
This paper focuses on the decentralized stochastic optimization problem f(x)=1/m ∑i=1m fi(x) over a connected network of n agents, where each local function has the form of fi(x) = E[F(x;ξ i)] which satisfies the (L0,L1)-smooth condition but possibly nonconvex and each random variable ξi follows distribution Di. We propose a novel algorithm called decentralized normalized stochastic gradient descent (DNSGD), which can achieve an ε-stationary point at each local agent. We present a new framework for analyzing decentralized first-order methods in the (L0,L1)-smooth setting, based on the Lyapunov function related to the product of the gradient norm and the consensus error. We show that the proposed algorithm attains the upper bounds on the sample complexity of O(m-1 (Lfσ2Δfε-4 + σ2ε-2 + Lf-2 L13σ2?fΔ-1 + Lf-2 L12σ2)) per agent and the communication complexity of O ((Lfε-2 + L1ε-1)γ-1/2 Δf), where Lf=L0 +L1ζ, σ2 is the variance of the stochastic gradient, Δf is the initial optimal function value gap, γ is the spectral gap of the network, and ζ is the degree of the gradient dissimilarity. In the special case of L1=0, the above results (nearly) match the lower bounds of decentralized stochastic nonconvex optimization under the standard smoothness. We also conduct numerical experiments to show the empirical superiority of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3dd4e66c-f297-4ac9-9bc5-5df10c0ad2f9Builds on23
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 139 citations
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 132 citations
Related papers
- Decentralized Riemannian Gradient Descent on the Stiefel ManifoldShixiang Chen, Alfredo García, Mingyi Hong, Shahin ShahrampourICML 2021 · 64 citations
- On the Complexity of Finite-Sum Smooth Optimization under the Polyak-Łojasiewicz ConditionYunyan Bai, Yuxing Liu, Luo LuoICML 2024 · 2 citations
- Topology-aware Generalization of Decentralized SGDTongtian Zhu, Fengxiang He, Lan Zhang, Zhengyang Niu et al.ICML 2022 · 58 citations
- Efficient Decentralized Stochastic Gradient Descent Method for Nonconvex Finite-Sum Optimization ProblemsWenkang Zhan, Gang Wu, Hongchang GaoAAAI 2022 · 8 citations
- A Hybrid Variance-Reduced Method for Decentralized Stochastic Non-Convex OptimizationRan Xin, Usman A. Khan, Soummya KarICML 2021 · 51 citations
