Lune

ICML2026Top-tier venue

From Optimization to Generalization under Heavy-Tailed Data: The Role of Gradient Clipping

Aleksandr Shestakov, Martin Takac, Eduard Gorbunov

2026Year

Abstract

Gradient clipping is widely used to stabilize stochastic gradient methods and is often theoretically motivated by heavy-tailed gradient noise, where even second moments may be infinite, seemingly contradicting the finite-sum ERM setting, where all empirical moments are finite once the dataset is fixed. We resolve this paradox by explicitly separating data sampling from optimization randomness: although moments are finite conditional on the dataset, heavy-tailed data induce dataset-dependent noise whose second moment typically grows with the dataset size NN. In particular, when ∥∇f(x⋆,ξ)∥\|\nabla f(x_\star,\xi)\| has tail index α∈(1,2)\alpha \in (1,2), the quantity 1N∑i=1N∥∇f(x⋆,ξi)∥2\frac{1}{N}\sum_{i=1}^N\|\nabla f(x_\star,\xi_i)\|^2 scales as N2α−1N^{\frac{2}{\alpha}-1}, leading to deteriorating convergence guarantees for standard SGD as NN increases. In contrast, we show that SGD with clipping avoids this growth and admits finite-sum convergence guarantees under heavy-tailed data for broad step-size and clipping schedules. We further derive generalization bounds for strongly convex smooth objectives and show that the tail behavior of gradients at the population minimizer is the key quantity linking optimization and generalization under heavy-tailed data.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines