TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent Kernels
Yaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma, Michael I. Jordan
摘要
State-of-the-art federated learning methods can perform far worse than their centralized counterparts when clients have dissimilar data distributions. For neural networks, even when centralized SGD easily finds a solution that is simultaneously performant for all clients, current federated optimization methods fail to converge to a comparable solution. We show that this performance disparity can largely be attributed to optimization challenges presented by nonconvexity. Specifically, we find that the early layers of the network do learn useful features, but the final layers fail to make use of them. That is, federated optimization applied to this non-convex problem distorts the learning of the final layers. Leveraging this observation, we propose a Train-Convexify-Train (TCT) procedure to sidestep this issue: first, learn features using off-the-shelf methods (e.g., FedAvg); then, optimize a convexified problem obtained from the network's empirical neural tangent kernel approximation. Our technique yields accuracy improvements of up to +36% on FMNIST and +37% on CIFAR10 when clients have dissimilar data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- Data Acquisition via Experimental Design for Data MarketsCharles Lu, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma 等NeurIPS 2024 · 被引用 12 次
- Global Convergence Analysis of Local SGD for Two-layer Neural Network without OverparameterizationYajie Bao, Amarda Shehu, Mingrui LiuNeurIPS 2023 · 被引用 8 次
- FedExP: Speeding Up Federated Averaging via ExtrapolationDivyansh Jhunjhunwala, Shiqiang Wang, Gauri JoshiICLR 2023 · 被引用 8 次
- Breaking Data Silos in Parkinson's Disease Diagnosis: An Adaptive Federated Learning Approach for Privacy-Preserving Facial Expression AnalysisMeng Pang, Houwei Xu, Zheng Huang, Yintao Zhou 等AAAI 2025 · 被引用 6 次
它引用的顶会 Paper30
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone 等CCS 2017 · 被引用 3,936 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi 等NeurIPS 2020 · 被引用 2,231 次
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 被引用 1,615 次
相关 Paper
- FL-NTK: A Neural Tangent Kernel-based Framework for Federated Learning AnalysisBaihe Huang, Xiaoxiao Li, Zhao Song, Xin YangICML 2021 · 被引用 66 次
- NTK-DFL: Enhancing Decentralized Federated Learning in Heterogeneous Settings via Neural Tangent KernelGabriel Thompson, Kai Yue, Chau-Wai Wong, Huaiyu DaiICML 2025
- FedAvg Converges to Zero Training Loss Linearly for Overparameterized Multi-Layer Neural NetworksBingqing Song, Prashant Khanduri, Xinwei Zhang, Jinfeng Yi 等ICML 2023 · 被引用 10 次
- On the Effectiveness of Partial Variance Reduction in Federated Learning with Heterogeneous DataBo Li, Mikkel N. Schmidt, Tommy S. Alstrøm, Sebastian U. StichCVPR 2023
- Heterogeneous Personalized Federated Learning by Local-Global Updates Mixing via Convergence RateMeirui Jiang, Anjie Le, Xiaoxiao Li, Qi DouICLR 2024 · 被引用 13 次
