Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction
Wei Jiang, Sifan Yang, Wenhao Yang, Lijun Zhang
摘要
Sign stochastic gradient descent (signSGD) is a communication-efficient method that transmits only the sign of stochastic gradients for parameter updating. Existing literature has demonstrated that signSGD can achieve a convergence rate of , where represents the dimension and is the iteration number. In this paper, we improve this convergence rate to by introducing the Sign-based Stochastic Variance Reduction (SSVR) method, which employs variance reduction estimators to track gradients and leverages their signs to update. For finite-sum problems, our method can be further enhanced to achieve a convergence rate of , where denotes the number of component functions. Furthermore, we investigate the heterogeneous majority vote in distributed settings and introduce two novel algorithms that attain improved convergence rates of and respectively, outperforming the previous results of and , where represents the number of nodes. Numerical experiments across different tasks validate the effectiveness of our proposed methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 被引用 37 次
- Adaptive Variance Reduction for Stochastic Optimization under Weaker AssumptionsWei Jiang, Sifan Yang, Yibo Wang, Lijun ZhangNeurIPS 2024 · 被引用 11 次
- Mirror Descent Under Generalized SmoothnessDingzhi Yu, Wei Jiang, Hongyi Tao, Yuanyu Wan 等ICML 2026 · 被引用 9 次
- Convergence Analysis of the Lion Optimizer in Centralized and Distributed SettingsWei Jiang, Mao Xu, Wenhao Yang, Yibo Wang 等ICML 2026 · 被引用 6 次
- SERENA: A Unified Stochastic Recursive Variance Reduced Gradient Framework for Riemannian Non-Convex OptimizationYan Liu, Mingjie Chen, Chaojie Ji, Hao Zhang 等ICML 2025
它引用的顶会 Paper12
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 被引用 164 次
- Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based AlgorithmsXiangyi Chen, Tiancong Chen, Haoran Sun, Zhiwei Steven Wu 等NeurIPS 2020 · 被引用 90 次
- Stochastic Sign Descent Methods: New Algorithms and Better TheoryMher Safaryan, Peter RichtárikICML 2021 · 被引用 70 次
- Generalized-Smooth Nonconvex Optimization is As Efficient As Smooth Nonconvex OptimizationZiyi Chen, Yi Zhou, Yingbin Liang, Zhaosong LuICML 2023 · 被引用 58 次
- Momentum Ensures Convergence of SIGNSGD under Weaker AssumptionsTao Sun, Qingsong Wang, Dongsheng Li, Bao WangICML 2023 · 被引用 36 次
相关 Paper
- Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced TopologyYan Huang, Ying Sun, Zehan Zhu, Changzhi Yan 等ICML 2022 · 被引用 18 次
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 被引用 131 次
- Asynchronous Decentralized Optimization With Implicit Stochastic Variance ReductionKenta Niwa, Guoqiang Zhang, W. Bastiaan Kleijn, Noboru Harada 等ICML 2021 · 被引用 16 次
- Stochastic Distributed Optimization under Average Second-order Similarity: Algorithms and AnalysisDachao Lin, Yuze Han, Haishan Ye, Zhihua ZhangNeurIPS 2023 · 被引用 17 次
- Bias-Variance Reduced Local SGD for Less Heterogeneous Federated LearningTomoya Murata, Taiji SuzukiICML 2021 · 被引用 61 次
