Optimizing Split Federated Learning through Adaptive Pipeline Parallelism
Zuan Xie, Yang Xu, Yunming Liao, Junhao Cheng, Jingjing Yang, Ying Zhu
摘要
Split federated learning (SFL) offers a promising solution for training large models on resource-constrained edge devices. However, existing synchronous and asynchronous SFL frameworks are plagued by a fundamental sequential execution pattern for communication and computation. This design creates a severe communication bottleneck, particularly for modern large models with substantial activation sizes, leading to significant resource underutilization and extended training time. To dismantle this bottleneck, we propose MicroSFL, a novel SFL framework with adaptive pipeline parallelism. MicroSFL enables workers to divide each of their training mini-batches into multiple small micro-batches and process them in a pipelined fashion. To counteract the challenges of system and statistical heterogeneity, MicroSFL adaptively assigns distinct micro-batch sizes to workers to accommodate their heterogeneous capacities, maximizing the temporal overlap between the computation and transmission of activations/gradients. Besides, to address statistical heterogeneity, MicroSFL employs an adaptive updating strategy. This strategy accumulates gradients at the server side to simulate updates on a large, balanced batch. Guided by a theoretical convergence analysis, the optimal updating weights are assigned to different workers to normalize their contributions for model updating. MicroSFL then dynamically determines the appropriate timing to apply the aggregated gradients, ensuring both stable convergence and high model accuracy. Extensive evaluations on a physical platform of 80 edge devices show that MicroSFL accelerates training by 3.64× to 5.67× and improves model accuracy by 3.6% to 6.1% compared to the baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ParallelSFL: A Novel Split Federated Learning Framework Tackling Heterogeneity IssuesYunming Liao, Yang Xu, Hongli Xu, Zhiwei Yao 等MobiCom 2024 · 被引用 27 次
- MergeSFL: Split Federated Learning with Feature Merging and Batch Size RegulationYunming Liao, Yang Xu, Hongli Xu, Lun Wang 等ICDE 2024 · 被引用 41 次
- FedASMU: Efficient Asynchronous Federated Learning with Dynamic Staleness-Aware Model UpdateJi Liu, Juncheng Jia, Tianshi Che, Chao Huo 等AAAI 2024 · 被引用 87 次
- Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update ApproachDandan Liang, Jianing Zhang, Evan Chen, Zhe Li 等NeurIPS 2025 · 被引用 8 次
- No One Idles: Efficient Heterogeneous Federated Learning with Parallel Edge and Server ComputationFeilong Zhang, Xianming Liu, Shiyi Lin, Gang Wu 等ICML 2023 · 被引用 15 次
