Runtime Variation in Big Data Analytics
Yiwen Zhu, Rathijit Sen, Robert Horton, John Mark Agosta
摘要
The dynamic nature of resource allocation and runtime conditions on Cloud can result in high variability in a job's runtime across multiple iterations, leading to a poor experience. Identifying the sources of such variation and being able to predict and adjust for them is crucial to cloud service providers to design reliable data processing pipelines, provision and allocate resources, adjust pricing services, meet SLOs and debug performance hazards. In this paper, we analyze the runtime variation of millions of production Scope jobs on Cosmos, an exabyte-scale internal analytics platform at Microsoft. We propose an innovative 2-step approach to predict job runtime distribution by characterizing typical distribution shapes combined with a classification model with an average accuracy of >96%, using an innovative interpretable machine-learning algorithm out-performing traditional regression models and better capturing long tails. We examine factors such as job plan characteristics and inputs, resource allocation, physical cluster heterogeneity and utilization, and scheduling policies. To the best of our knowledge, this is the first study on predicting categories of runtime distributions for enterprise analytics workloads at scale. Furthermore, we examine how our methods can be used to analyze what-if scenarios, focusing on the impact of resource allocation, scheduling, and physical cluster provisioning decisions on a job's runtime consistency and predictability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- A Statistical Perspective on Discovering Functional Dependencies in Noisy DataYunjia Zhang, Zhihan Guo, Theodoros RekatsinasSIGMOD 2020 · 被引用 45 次
- Phoebe: A Learning-based Checkpoint OptimizerYiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das 等VLDB 2021 · 被引用 10 次
- Unearthing inter-job dependencies for better cluster schedulingAndrew Chung, Subru Krishnan, Konstantinos Karanasos, Carlo Curino 等OSDI 2020 · 被引用 8 次
- ReLoca: Optimize Resource Allocation for Data-parallel Jobs using Deep LearningZhiyao Hu, Dongsheng Li, Dongxiang Zhang, Yixin ChenINFOCOM 2020 · 被引用 3 次
相关 Paper
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- Unlocking unallocated cloud capacity for long, uninterruptible workloadsAnup Agarwal, Shadi A. Noghabi, Íñigo Goiri, Srinivasan Seshan 等NSDI 2023 · 被引用 9 次
- A Case for Task Sampling based Learning for Cluster Job SchedulingAkshay Jajoo, Y. Charlie Hu, Xiaojun Lin, Nan DengNSDI 2022 · 被引用 35 次
- Generating Complex, Realistic Cloud Workloads using Recurrent Neural NetworksShane Bergsma, Timothy Zeyl, Arik Senderovich, J. Christopher BeckSOSP 2021 · 被引用 21 次
- Caerus: NIMBLE Task Scheduling for Serverless AnalyticsHong Zhang, Yupeng Tang, Anurag Khandelwal, Jingrong Chen 等NSDI 2021 · 被引用 75 次
