Towards Demystifying Serverless Machine Learning Training
Jiawei Jiang, Shaoduo Gan, Yue Liu, Fanlin Wang, Gustavo Alonso, Ana Klimovic, Ankit Singla, Wentao Wu, Ce Zhang
摘要
The appeal of serverless (FaaS) has triggered a growing interest on how to use it in data-intensive applications such as ETL, query processing, or machine learning (ML). Several systems exist for training large-scale ML models on top of serverless infrastructures (e.g., AWS Lambda) but with inconclusive results in terms of their performance and relative advantage over "serverful" infrastructures (IaaS). In this paper we present a systematic, comparative study of distributed ML training over FaaS and IaaS. We present a design space covering design choices such as optimization algorithms and synchronization protocols, and implement a platform, LambdaML, that enables a fair comparison between FaaS and IaaS. We present experimental results using LambdaML, and further develop an analytic model to capture cost/performance tradeoffs that must be considered when opting for a serverless infrastructure. Our results indicate that ML training pays off in serverless only for models with efficient (i.e., reduced) communication and that quickly converge. In general, FaaS can be much faster but it is never significantly cheaper than IaaS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- FaaSNet: Scalable and Fast Provisioning of Custom Serverless Container Runtimes at Alibaba Cloud Function ComputeAo Wang, Shuai Chang, Huangshi Tian, Hongqi Wang 等USENIX ATC 2021 · 被引用 171 次
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 被引用 76 次
- Palette Load Balancing: Locality Hints for Serverless FunctionsMania Abdi, Samuel Ginzburg, Xiayue Charles Lin, Jose M. Faleiro 等EuroSys 2023 · 被引用 44 次
- MXFaaS: Resource Sharing in Serverless Environments for Parallelism and EfficiencyJovan Stojkovic, Tianyin Xu, Hubertus Franke, Josep TorrellasISCA 2023 · 被引用 39 次
- EAVS: Edge-assisted Adaptive Video Streaming with Fine-grained Serverless PipelinesBiao Hou, Song Yang, Fernando A. Kuipers, Lei Jiao 等INFOCOM 2023 · 被引用 26 次
它引用的顶会 Paper5
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud InfrastructureIngo Müller, Renato Marroquín, Gustavo AlonsoSIGMOD 2020 · 被引用 135 次
- KungFu: Making Training in Distributed Machine Learning AdaptiveLuo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis 等OSDI 2020 · 被引用 92 次
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm 等SIGMOD 2020 · 被引用 27 次
- Dynamic Parameter Allocation in Parameter ServersAlexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker MarklVLDB 2020 · 被引用 18 次
相关 Paper
- FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud CommunicationJoe Oakley, Hakan FerhatosmanogluICDE 2024 · 被引用 6 次
- Serverless Data Science - Are We There Yet? A Case Study of Model ServingYuncheng Wu, Tien Tuan Anh Dinh, Guoyu Hu, Meihui Zhang 等SIGMOD 2022 · 被引用 27 次
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 被引用 184 次
- Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless ComputingMengfan Liu, Wei Wang, Chuan WuINFOCOM 2025 · 被引用 6 次
- Faasm: Lightweight Isolation for Efficient Stateful Serverless ComputingSimon Shillaker, Peter R. PietzuchUSENIX ATC 2020 · 被引用 382 次
