Lune

ICML2023Top-tier venue

Momentum Ensures Convergence of SIGNSGD under Weaker Assumptions

Tao Sun, Qingsong Wang, Dongsheng Li, Bao Wang

2023Year
36Citations
13Top-tier citations

Abstract

Sign Stochastic Gradient Descent (SIGNSGD) is a communication-efficient stochastic algorithm that only uses the sign information of the stochastic gradient to update the model's weights. However, the existing convergence theory of SIGNSGD either requires increasing batch sizes during training or assumes the gradient noise is symmetric and unimodal. Error feedback has been used to guarantee the convergence of SIGNSGD under weaker assumptions at the cost of communication overhead. This paper revisits the convergence of SIGNSGD and proves that momentum can remedy SIGNSGD under weaker assumptions than previous techniques; in particular, our convergence theory does not require the assumption of bounded stochastic gradient or increased batch size. Our results resonate with echoes of previous empirical results where, unlike SIGNSGD, SIGNSGD with momentum maintains good performance even with small batch sizes. Another new result is that SIGNSGD with momentum can achieve an improved convergence rate when the objective function is second-order smooth. We further extend our theory to SIGNSGD with major vote and federated learning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers13

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines