KungFu: Making Training in Distributed Machine Learning Adaptive
Luo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis, Andrei-Octavian Brabete, Peter R. Pietzuch
摘要
When using distributed machine learning (ML) systems to train models on a cluster of worker machines, users must con-figure a large number of parameters: hyper-parameters (e.g. the batch size and the learning rate) affect model convergence; system parameters (e.g. the number of workers and their communication topology) impact training performance. In current systems, adapting such parameters during training is ill-supported. Users must set system parameters at deployment time, and provide fixed adaptation schedules for hyper-parameters in the training program. We describe Kung Fu, a distributed ML library for Tensor-Flow that is designed to enable adaptive training. Kung Fu allows users to express high-level Adaptation Policies(APs)that describe how to change hyper- and system parameters during training. APs take real-time monitored metrics (e.g. signal-to-noise ratios and noise scale) as input and trigger control actions (e.g. cluster rescaling or synchronisation strategy updates). For execution, APs are translated into monitoring and control operators, which are embedded in the data flowgraph. APs exploit an efficient asynchronous collective communication layer, which ensures concurrency and consistency of monitoring and adaptation operations
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Towards Demystifying Serverless Machine Learning TrainingJiawei Jiang, Shaoduo Gan, Yue Liu, Fanlin Wang 等SIGMOD 2021 · 被引用 107 次
- ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningDiandian Gu, Yihao Zhao, Yinmin Zhong, Yifan Xiong 等ASPLOS 2023 · 被引用 62 次
- Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine LearningPengfei Zheng, Rui Pan, Tarannum Khan, Shivaram Venkataraman 等NSDI 2023 · 被引用 56 次
- Bringing Decentralized Search to Decentralized ServicesMingyu Li, Jinhao Zhu, Tianxu Zhang, Cheng Tan 等OSDI 2021 · 被引用 28 次
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun 等SC 2023 · 被引用 16 次
它引用的顶会 Paper2
相关 Paper
- Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DMLJinkun Geng, Dan Li, Shuai WangICDE 2020 · 被引用 5 次
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams 等SC 2020 · 被引用 21 次
- BAGUA: Scaling up Distributed Learning with System RelaxationsShaoduo Gan, Xiangru Lian, Rui Wang, Jianbin Chang 等VLDB 2022 · 被引用 35 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- Modeling and Optimizing the Scaling Performance in Distributed Deep Learning TrainingTing Liu, Tianhao Miao, Qinghua Wu, Zhenyu Li 等WWW 2022 · 被引用 7 次
