Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction
Wei Jiang, Sifan Yang, Wenhao Yang, Lijun Zhang
Abstract
Sign stochastic gradient descent (signSGD) is a communication-efficient method that transmits only the sign of stochastic gradients for parameter updating. Existing literature has demonstrated that signSGD can achieve a convergence rate of , where represents the dimension and is the iteration number. In this paper, we improve this convergence rate to by introducing the Sign-based Stochastic Variance Reduction (SSVR) method, which employs variance reduction estimators to track gradients and leverages their signs to update. For finite-sum problems, our method can be further enhanced to achieve a convergence rate of , where denotes the number of component functions. Furthermore, we investigate the heterogeneous majority vote in distributed settings and introduce two novel algorithms that attain improved convergence rates of and respectively, outperforming the previous results of and , where represents the number of nodes. Numerical experiments across different tasks validate the effectiveness of our proposed methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6b07da1-6d0d-4b21-aebd-35231ca093ecCited by top-tier papers6
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 37 citations
- Adaptive Variance Reduction for Stochastic Optimization under Weaker AssumptionsWei Jiang, Sifan Yang, Yibo Wang, Lijun ZhangNeurIPS 2024 · 11 citations
- Mirror Descent Under Generalized SmoothnessDingzhi Yu, Wei Jiang, Hongyi Tao, Yuanyu Wan et al.ICML 2026 · 9 citations
- Convergence Analysis of the Lion Optimizer in Centralized and Distributed SettingsWei Jiang, Mao Xu, Wenhao Yang, Yibo Wang et al.ICML 2026 · 6 citations
- SERENA: A Unified Stochastic Recursive Variance Reduced Gradient Framework for Riemannian Non-Convex OptimizationYan Liu, Mingjie Chen, Chaojie Ji, Hao Zhang et al.ICML 2025
Builds on12
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 164 citations
- Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based AlgorithmsXiangyi Chen, Tiancong Chen, Haoran Sun, Zhiwei Steven Wu et al.NeurIPS 2020 · 90 citations
- Stochastic Sign Descent Methods: New Algorithms and Better TheoryMher Safaryan, Peter RichtárikICML 2021 · 70 citations
- Generalized-Smooth Nonconvex Optimization is As Efficient As Smooth Nonconvex OptimizationZiyi Chen, Yi Zhou, Yingbin Liang, Zhaosong LuICML 2023 · 58 citations
- Momentum Ensures Convergence of SIGNSGD under Weaker AssumptionsTao Sun, Qingsong Wang, Dongsheng Li, Bao WangICML 2023 · 36 citations
Related papers
- Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced TopologyYan Huang, Ying Sun, Zehan Zhu, Changzhi Yan et al.ICML 2022 · 18 citations
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- Asynchronous Decentralized Optimization With Implicit Stochastic Variance ReductionKenta Niwa, Guoqiang Zhang, W. Bastiaan Kleijn, Noboru Harada et al.ICML 2021 · 16 citations
- Stochastic Distributed Optimization under Average Second-order Similarity: Algorithms and AnalysisDachao Lin, Yuze Han, Haishan Ye, Zhihua ZhangNeurIPS 2023 · 17 citations
- Bias-Variance Reduced Local SGD for Less Heterogeneous Federated LearningTomoya Murata, Taiji SuzukiICML 2021 · 61 citations
