Lune

ICML2026顶会

From Optimization to Generalization under Heavy-Tailed Data: The Role of Gradient Clipping

Aleksandr Shestakov, Martin Takac, Eduard Gorbunov

出版方
2026年份

摘要

Gradient clipping is widely used to stabilize stochastic gradient methods and is often theoretically motivated by heavy-tailed gradient noise, where even second moments may be infinite, seemingly contradicting the finite-sum ERM setting, where all empirical moments are finite once the dataset is fixed. We resolve this paradox by explicitly separating data sampling from optimization randomness: although moments are finite conditional on the dataset, heavy-tailed data induce dataset-dependent noise whose second moment typically grows with the dataset size NN. In particular, when ∥∇f(x⋆,ξ)∥\|\nabla f(x_\star,\xi)\| has tail index α∈(1,2)\alpha \in (1,2), the quantity 1N∑i=1N∥∇f(x⋆,ξi)∥2\frac{1}{N}\sum_{i=1}^N\|\nabla f(x_\star,\xi_i)\|^2 scales as N2α−1N^{\frac{2}{\alpha}-1}, leading to deteriorating convergence guarantees for standard SGD as NN increases. In contrast, we show that SGD with clipping avoids this growth and admits finite-sum convergence guarantees under heavy-tailed data for broad step-size and clipping schedules. We further derive generalization bounds for strongly convex smooth objectives and show that the tail behavior of gradients at the population minimizer is the key quantity linking optimization and generalization under heavy-tailed data.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖