Improved Convergence in High Probability of Clipped Gradient Methods with Heavy Tailed Noise
Ta Duy Nguyen, Thien Hang Nguyen, Alina Ene, Huy L. Nguyen
摘要
In this work, we study the convergence in high probability of clipped gradient 1 methods when the noise distribution has heavy tails, i.e., with bounded p th mo-2 ments, for some 1 < p ≤ 2 . Prior works in this setting follow the same recipe of 3 using concentration inequalities and an inductive argument with union bound to 4 bound the iterates across all iterations. This method results in an increase in the 5 failure probability by a factor of T , where T is the number of iterations. We in-6 stead propose a new analysis approach based on bounding the moment generating 7 function of a well chosen supermartingale sequence. We improve the dependency 8 on T in the convergence guarantee for a wide range of algorithms with clipped 9 gradients, including stochastic (accelerated) mirror descent for convex objectives 10 and stochastic gradient descent for nonconvex objectives. Our high probability 11 bounds achieve the optimal convergence rates and match the best currently known 12 in-expectation bounds. Our approach naturally allows the algorithms to use time-13 varying step sizes and clipping parameters when the time horizon is unknown, 14 which appears difficult or even impossible using the techniques from prior works. 15 Furthermore, we show that in the case of clipped stochastic mirror descent, several 16 problem constants, including the initial distance to the optimum, are not required 17 when setting step sizes and clipping parameters. 18
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Revisiting the Last-Iterate Convergence of Stochastic Gradient MethodsZijian Liu, Zhengyuan ZhouICLR 2024 · 被引用 32 次
- A Huber Loss Minimization Approach to Mean Estimation under User-level Differential PrivacyPuning Zhao, Lifeng Lai, Li Shen, Qingming Li 等NeurIPS 2024 · 被引用 17 次
- Understanding Stochastic Natural Gradient Variational InferenceKaiwen Wu, Jacob R. GardnerICML 2024 · 被引用 11 次
- High-Probability Bound for Non-Smooth Non-Convex Stochastic Optimization with Heavy TailsLangqi Liu, Yibo Wang, Lijun ZhangICML 2024 · 被引用 11 次
- Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal ConvergenceShuhua Yu, Dusan Jakovetic, Soummya KarICLR 2026 · 被引用 7 次
它引用的顶会 Paper10
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient ClippingEduard Gorbunov, Marina Danilova, Alexander V. GasnikovNeurIPS 2020 · 被引用 181 次
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy TailsAshok Cutkosky, Harsh MehtaNeurIPS 2021 · 被引用 119 次
- High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded VarianceAbdurakhmon Sadiev, Marina Danilova, Eduard Gorbunov, Samuel Horváth 等ICML 2023 · 被引用 68 次
相关 Paper
- High-Probability Bounds for the Last Iterate of Clipped SGDSavelii Chezhegov, Daniela Angela Parletta, Andrea Paudice, Eduard GorbunovICLR 2026
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined AnalysisZijian LiuICLR 2026 · 被引用 5 次
- Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient ClippingZijian Liu, Zhengyuan ZhouICLR 2025
- Stochastic Gradient Methods under Heavy-Tailed Noises in Weakly Convex OptimizationTianxi Zhu, Yi Xu, Qi Wang, Xiangyang JiICML 2026
- High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy TailsShaojie Li, Yong LiuICML 2022 · 被引用 37 次
