SoCFlow: Efficient and Scalable DNN Training on SoC-Clustered Edge Servers
Daliang Xu, Mengwei Xu, Chiheng Lou, Li Zhang, Gang Huang, Xin Jin, Xuanzhe Liu
Abstract
SoC-Cluster, a novel server architecture composed of massive mobile system-on-chips (SoCs), is gaining popularity in industrial edge computing due to its energy efficiency and compatibility with existing mobile applications. However, we observe that the deployed SoC-Cluster servers are not fully utilized, because the hosted workloads are mostly usertriggered and have significant tidal phenomena. To harvest the free cycles, we propose to co-locate deep learning tasks on them.
We present SoCFlow, the first framework that can efficiently train deep learning models on SoC-Cluster. To deal with the intrinsic inadequacy of commercial SoC-Cluster servers, SoCFlow incorporates two novel techniques: (1) the group-wise parallelism with delayed aggregation that can train deep learning models fast and scalably without being influenced by the network bottleneck; (2) the data-parallel mixed-precision training algorithm that can fully unleash the heterogeneous processors' capability of mobile SoCs. We have fully implemented SoCFlow and demonstrated its effectiveness through extensive experiments. The experiments show that SoCFlow significantly and consistently outperforms all baselines regarding the training speed while preserving the convergence accuracy, e.g., 1.6×-740× convergence speedup with 32 SoCs. Compared to commodity GPU (NVIDIA V100) under the same power budget, SoCFlow * : Equal contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5029bd5f-b5f0-4e66-9335-ff38d97c5755Cited by top-tier papers1
Ask how each one uses itBuilds on17
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Billion-scale federated learning on mobile clients: a submodel design with tunable privacyChaoyue Niu, Fan Wu, Shaojie Tang, Lifeng Hua et al.MobiCom 2020 · 114 citations
- Distribution Adaptive INT8 Quantization for Training CNNsKang Zhao, Sida Huang, Pan Pan, Yinghan Li et al.AAAI 2021 · 86 citations
Related papers
- More is Different: Prototyping and Analyzing a New Form of Edge Server with Massive Mobile SoCsLi Zhang, Zhe Fu, Boqing Shi, Xiang Li et al.USENIX ATC 2024
- ParallelSFL: A Novel Split Federated Learning Framework Tackling Heterogeneity IssuesYunming Liao, Yang Xu, Hongli Xu, Zhiwei Yao et al.MobiCom 2024 · 27 citations
- High-density Mobile Cloud Gaming on Edge SoC ClustersLi Zhang, Shangguang Wang, Mengwei XuUSENIX ATC 2024 · 6 citations
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
- Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU ClustersZiyue Luo, Jia Liu, Myungjin Lee, Ness B. ShroffINFOCOM 2025 · 5 citations
