Taming Fat-Tailed ("Heavier-Tailed" with Potentially Infinite Variance) Noise in Federated Learning
Haibo Yang, Peiwen Qiu, Jia Liu
Abstract
In recent years, federated learning (FL) has emerged as an important distributed machine learning paradigm to collaboratively learn a global model with multiple clients, while keeping data local and private. However, a key assumption in most existing works on FL algorithms' convergence analysis is that the noise in stochastic first-order information has a finite variance. Although this assumption covers all light-tailed (i.e., sub-exponential) and some heavy-tailed noise distributions (e.g., log-normal, Weibull, and some Pareto distributions), it fails for many fat-tailed noise distributions (i.e., "heavier-tailed" with potentially infinite variance) that have been empirically observed in the FL literature. To date, it remains unclear whether one can design convergent algorithms for FL systems that experience fat-tailed noise. This motivates us to fill this gap in this paper by proposing an algorithmic framework called FAT-Clipping (federated averaging with two-sided learning rates and clipping), which contains two variants: FAT-Clipping per-round (FAT-Clipping-PR) and FAT-Clipping per-iteration (FAT-Clipping-PI). Specifically, for the largest tail-index α ∈ (1, 2] such that the fat-tailed noise in FL still has a bounded α-moment, we show that both variants achieve O((mT ) 2-α α ) and O((mT ) 1-α 3α-2 ) convergence rates in the strongly-convex and general non-convex settings, respectively, where m and T are the numbers of clients and communication rounds. Moreover, with more clipping operations compared to FAT-Clipping-PR, FAT-Clipping-PI further enjoys a linear speedup effect with respect to the number of local updates at each client and being lower-bound-matching (i.e., order-optimal). Collectively, our results advance the understanding of designing efficient algorithms for FL systems that exhibit fat-tailed first-order oracle information. 2 Related work In this section, we will provide a quick overview on three related topics in the literature: i) federated learning, ii) heavy-tailed noise in learning, and iii) the clipping techniques, thus putting our work into comparative perspective to highlight our novelty and differences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Federated Multi-Objective LearningHaibo Yang, Zhuqing Liu, Jia Liu, Chaosheng Dong et al.NeurIPS 2023 · 28 citations
- An improved analysis of per-sample and per-update clipping in federated learningBo Li, Xiaowen Jiang, Mikkel N. Schmidt, Tommy Sonne Alstrøm et al.ICLR 2024 · 9 citations
- Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal ConvergenceShuhua Yu, Dusan Jakovetic, Soummya KarICLR 2026 · 7 citations
- Tight High-Probability Bounds for Nonconvex Heavy-Tailed Scenario under Weaker AssumptionsWeixin An, Yuanyuan Liu, Fanhua Shang, Han Yu et al.NeurIPS 2025
- Efficient Distributed Optimization under Heavy-Tailed NoiseSu Hyeong Lee, Manzil Zaheer, Tian LiICML 2025
Builds on24
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi et al.NeurIPS 2020 · 2,231 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
Related papers
- Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient ClippingZijian Liu, Zhengyuan ZhouICLR 2025
- Understanding Clipping for Federated Learning: Convergence and Client-Level Differential PrivacyXinwei Zhang, Xiangyi Chen, Mingyi Hong, Steven Wu et al.ICML 2022 · 134 citations
- High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed NoiseEduard Gorbunov, Abdurakhmon Sadiev, Marina Danilova, Samuel Horváth et al.ICML 2024 · 27 citations
- Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGradZijian LiuICML 2026 · 3 citations
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined AnalysisZijian LiuICLR 2026 · 5 citations
