FSL-SAGE: Accelerating Federated Split Learning via Smashed Activation Gradient Estimation
Srijith Nair, Michael Lin, Peizhong Ju, Amirreza Talebi, Elizabeth Serena Bentley, Jia Liu
摘要
Collaborative training methods like Federated Learning (FL) and Split Learning (SL) enable distributed machine learning without sharing raw data. However, FL assumes clients can train entire models, which is infeasible for large-scale models. In contrast, while SL alleviates the client memory constraint in FL by offloading most training to the server, it increases network latency due to its sequential nature. Other methods address the conundrum by using local loss functions for parallel client-side training to improve efficiency, but they lack server feedback and potentially suffer poor accuracy. We propose FSL-SAGE (Federated Split Learning via Smashed Activation Gradient Estimation), a new federated split learning algorithm that estimates server-side gradient feedback via auxiliary models. These auxiliary models periodically adapt to emulate server behavior on local datasets. We show that FSL-SAGE achieves a convergence rate of O(1/ √ T ), where T is the number of communication rounds. This result matches Fe-dAvg, while significantly reducing communication costs and client memory requirements. Our empirical results also verify that it outperforms existing state-of-the-art FSL methods, offering both communication efficiency and accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Group Knowledge Transfer: Federated Learning of Large CNNs at the EdgeChaoyang He, Murali Annavaram, Salman AvestimehrNeurIPS 2020 · 被引用 605 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated LearningHaibo Yang, Minghong Fang, Jia LiuICLR 2021 · 被引用 310 次
相关 Paper
- Convergence Analysis of Split Federated Learning on Heterogeneous DataPengchao Han, Chao Huang, Geng Tian, Ming Tang 等NeurIPS 2024 · 被引用 32 次
- LocFedMix-SL: Localize, Federate, and Mix for Improved Scalability, Convergence, and Latency in Split LearningSeungeun Oh, Jihong Park, Praneeth Vepakomma, Sihun Baek 等WWW 2022 · 被引用 66 次
- Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update ApproachDandan Liang, Jianing Zhang, Evan Chen, Zhe Li 等NeurIPS 2025 · 被引用 8 次
- Delayed Gradient Averaging: Tolerate the Communication Latency for Federated LearningLigeng Zhu, Hongzhou Lin, Yao Lu, Yujun Lin 等NeurIPS 2021 · 被引用 4 次
- AOCC-FL: Federated Learning with Aligned Overlapping via Calibrated CompensationHaozhao Wang, Wenchao Xu, Yunfeng Fan, Ruixuan Li 等INFOCOM 2023 · 被引用 8 次
