High-Probability Bound for Non-Smooth Non-Convex Stochastic Optimization with Heavy Tails
Langqi Liu, Yibo Wang, Lijun Zhang
Abstract
Recently, Cutkosky et al. introduce the onlineto-non-convex framework, which utilizes online learning methods to solve non-smooth nonconvex optimization problems, and achieves an O(ϵ -3 δ -1 ) gradient complexity for finding (δ, ϵ)stationary points. However, their results rely on the bounded variance assumption of stochastic gradients and only hold in expectation. To address these limitations, we investigate the case that stochastic gradients obey heavy-tailed distributions with finite p-th moments for some p ∈ (1, 2], and propose a novel algorithm which is able to identify a (δ, ϵ)-stationary point with high probability, after consuming Õ(ϵ -2p-1 p-1 δ -1 ) stochastic gradients. The key idea is first incorporating the gradient clipping technique into the onlineto-non-convex framework to produce a sequence of points, the averaged gradient norms of which is no greater than ϵ. Then, we propose a validation method to select one (δ, ϵ)-stationary point among the candidates. When gradient distributions have bounded variance, i.e., p = 2, our result turns into Õ(ϵ -3 δ -1 ), which improves the existing Õ(ϵ -4 δ -1 ) high-probability bound. When the objective is smooth, our algorithm can also find an ϵ-stationary point with Õ(ϵ -3p-2 p-1 ) gradient queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2b1a7cf-70f0-493b-8f94-99d8b1dca698Cited by top-tier papers5
- Mirror Descent Under Generalized SmoothnessDingzhi Yu, Wei Jiang, Hongyi Tao, Yuanyu Wan et al.ICML 2026 · 9 citations
- Nonlinearly Preconditioned Gradient Methods: Momentum and Stochastic AnalysisKonstantinos A. Oikonomidis, Jan Quan, Panagiotis PatrinosNeurIPS 2025 · 6 citations
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined AnalysisZijian LiuICLR 2026 · 5 citations
- Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient ClippingZijian Liu, Zhengyuan ZhouICLR 2025
- Stochastic Gradient Methods under Heavy-Tailed Noises in Weakly Convex OptimizationTianxi Zhu, Yi Xu, Qi Wang, Xiangyang JiICML 2026
Builds on16
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy TailsAshok Cutkosky, Harsh MehtaNeurIPS 2021 · 119 citations
- Gradient-Free Methods for Deterministic and Stochastic Nonsmooth Nonconvex OptimizationTianyi Lin, Zeyu Zheng, Michael I. JordanNeurIPS 2022 · 102 citations
Related papers
- Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double OptimismFrancisco Patitucci, Ruichen Jiang, Aryan MokhtariICLR 2026 · 3 citations
- Optimal Stochastic Non-smooth Non-convex Optimization through Online-to-Non-convex ConversionAshok Cutkosky, Harsh Mehta, Francesco OrabonaICML 2023 · 54 citations
- Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity LimitsAbdurakhmon Sadiev, Peter Richtárik, Ilyas FatkhullinNeurIPS 2025 · 4 citations
- Improved Convergence in High Probability of Clipped Gradient Methods with Heavy Tailed NoiseTa Duy Nguyen, Thien Hang Nguyen, Alina Ene, Huy L. NguyenNeurIPS 2023 · 65 citations
- High-Probability Bounds for the Last Iterate of Clipped SGDSavelii Chezhegov, Daniela Angela Parletta, Andrea Paudice, Eduard GorbunovICLR 2026
