Momentum Ensures Convergence of SIGNSGD under Weaker Assumptions
Tao Sun, Qingsong Wang, Dongsheng Li, Bao Wang
Abstract
Sign Stochastic Gradient Descent (SIGNSGD) is a communication-efficient stochastic algorithm that only uses the sign information of the stochastic gradient to update the model's weights. However, the existing convergence theory of SIGNSGD either requires increasing batch sizes during training or assumes the gradient noise is symmetric and unimodal. Error feedback has been used to guarantee the convergence of SIGNSGD under weaker assumptions at the cost of communication overhead. This paper revisits the convergence of SIGNSGD and proves that momentum can remedy SIGNSGD under weaker assumptions than previous techniques; in particular, our convergence theory does not require the assumption of bounded stochastic gradient or increased batch size. Our results resonate with echoes of previous empirical results where, unlike SIGNSGD, SIGNSGD with momentum maintains good performance even with small batch sizes. Another new result is that SIGNSGD with momentum can achieve an improved convergence rate when the objective function is second-order smooth. We further extend our theory to SIGNSGD with major vote and federated learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren et al.NeurIPS 2025 · 58 citations
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 37 citations
- Communication Efficient Distributed Training with Distributed LionBo Liu, Lemeng Wu, Lizhang Chen, Kaizhao Liang et al.NeurIPS 2024 · 21 citations
- Efficient Sign-Based Optimization: Accelerating Convergence via Variance ReductionWei Jiang, Sifan Yang, Wenhao Yang, Lijun ZhangNeurIPS 2024 · 19 citations
- Error Feedback for Muon and FriendsKaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan, Peter RichtárikICLR 2026 · 13 citations
Builds on6
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- Adam Can Converge Without Any Modification On Update RulesYushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun et al.NeurIPS 2022 · 134 citations
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang et al.NeurIPS 2022 · 111 citations
- Sign Bits Are All You Need for Black-Box AttacksAbdullah Al-Dujaili, Una-May O'ReillyICLR 2020 · 93 citations
- Stochastic Sign Descent Methods: New Algorithms and Better TheoryMher Safaryan, Peter RichtárikICML 2021 · 70 citations
Related papers
- SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign DecodingChanho Park, Namyoon LeeICML 2024 · 5 citations
- Towards Faster Decentralized Stochastic Optimization with Communication CompressionRustem Islamov, Yuan Gao, Sebastian U. StichICLR 2025
- Momentum Benefits Non-iid Federated Learning Simply and ProvablyZiheng Cheng, Xinmeng Huang, Pengfei Wu, Kun YuanICLR 2024 · 40 citations
- Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient NoiseRui Pan, Yuxing Liu, Xiaoyu Wang, Tong ZhangICLR 2024 · 10 citations
- Momentum Provably Improves Error Feedback!Ilyas Fatkhullin, Alexander Tyurin, Peter RichtárikNeurIPS 2023 · 47 citations
