FSL-SAGE: Accelerating Federated Split Learning via Smashed Activation Gradient Estimation
Srijith Nair, Michael Lin, Peizhong Ju, Amirreza Talebi, Elizabeth Serena Bentley, Jia Liu
Abstract
Collaborative training methods like Federated Learning (FL) and Split Learning (SL) enable distributed machine learning without sharing raw data. However, FL assumes clients can train entire models, which is infeasible for large-scale models. In contrast, while SL alleviates the client memory constraint in FL by offloading most training to the server, it increases network latency due to its sequential nature. Other methods address the conundrum by using local loss functions for parallel client-side training to improve efficiency, but they lack server feedback and potentially suffer poor accuracy. We propose FSL-SAGE (Federated Split Learning via Smashed Activation Gradient Estimation), a new federated split learning algorithm that estimates server-side gradient feedback via auxiliary models. These auxiliary models periodically adapt to emulate server behavior on local datasets. We show that FSL-SAGE achieves a convergence rate of O(1/ √ T ), where T is the number of communication rounds. This result matches Fe-dAvg, while significantly reducing communication costs and client memory requirements. Our empirical results also verify that it outperforms existing state-of-the-art FSL methods, offering both communication efficiency and accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74ca3987-6166-45a6-b19d-53715ff2ba19Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Group Knowledge Transfer: Federated Learning of Large CNNs at the EdgeChaoyang He, Murali Annavaram, Salman AvestimehrNeurIPS 2020 · 605 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated LearningHaibo Yang, Minghong Fang, Jia LiuICLR 2021 · 310 citations
Related papers
- Convergence Analysis of Split Federated Learning on Heterogeneous DataPengchao Han, Chao Huang, Geng Tian, Ming Tang et al.NeurIPS 2024 · 32 citations
- LocFedMix-SL: Localize, Federate, and Mix for Improved Scalability, Convergence, and Latency in Split LearningSeungeun Oh, Jihong Park, Praneeth Vepakomma, Sihun Baek et al.WWW 2022 · 66 citations
- Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update ApproachDandan Liang, Jianing Zhang, Evan Chen, Zhe Li et al.NeurIPS 2025 · 8 citations
- Delayed Gradient Averaging: Tolerate the Communication Latency for Federated LearningLigeng Zhu, Hongzhou Lin, Yao Lu, Yujun Lin et al.NeurIPS 2021 · 4 citations
- AOCC-FL: Federated Learning with Aligned Overlapping via Calibrated CompensationHaozhao Wang, Wenchao Xu, Yunfeng Fan, Ruixuan Li et al.INFOCOM 2023 · 8 citations
