Lune

NeurIPS2025Top-tier venue

Convergence of Clipped SGD on Convex (L0, L1)-Smooth Functions

Ofir Gaash, Kfir Y. Levy, Yair Carmon

2025Year
5Citations
5Top-tier citations

Abstract

We study stochastic gradient descent (SGD) with gradient clipping on convex functions under a generalized smoothness assumption called (L0,L1)(L_0,L_1)-smoothness. Using gradient clipping, we establish a high probability convergence rate that matches the SGD rate in the LL smooth case up to polylogarithmic factors and additive terms. We also propose a variation of adaptive SGD with gradient clipping, which achieves the same guarantee. We perform empirical experiments to examine our theory and algorithmic choices.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dc7664cd-1bd9-4ec5-a41c-08423d2256a5

Cited by top-tier papers5

Ask how each one uses it

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines