Runtime Variation in Big Data Analytics
Yiwen Zhu, Rathijit Sen, Robert Horton, John Mark Agosta
Abstract
The dynamic nature of resource allocation and runtime conditions on Cloud can result in high variability in a job's runtime across multiple iterations, leading to a poor experience. Identifying the sources of such variation and being able to predict and adjust for them is crucial to cloud service providers to design reliable data processing pipelines, provision and allocate resources, adjust pricing services, meet SLOs and debug performance hazards. In this paper, we analyze the runtime variation of millions of production Scope jobs on Cosmos, an exabyte-scale internal analytics platform at Microsoft. We propose an innovative 2-step approach to predict job runtime distribution by characterizing typical distribution shapes combined with a classification model with an average accuracy of >96%, using an innovative interpretable machine-learning algorithm out-performing traditional regression models and better capturing long tails. We examine factors such as job plan characteristics and inputs, resource allocation, physical cluster heterogeneity and utilization, and scheduling policies. To the best of our knowledge, this is the first study on predicting categories of runtime distributions for enterprise analytics workloads at scale. Furthermore, we examine how our methods can be used to analyze what-if scenarios, focusing on the impact of resource allocation, scheduling, and physical cluster provisioning decisions on a job's runtime consistency and predictability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- A Statistical Perspective on Discovering Functional Dependencies in Noisy DataYunjia Zhang, Zhihan Guo, Theodoros RekatsinasSIGMOD 2020 · 45 citations
- Phoebe: A Learning-based Checkpoint OptimizerYiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das et al.VLDB 2021 · 10 citations
- Unearthing inter-job dependencies for better cluster schedulingAndrew Chung, Subru Krishnan, Konstantinos Karanasos, Carlo Curino et al.OSDI 2020 · 8 citations
- ReLoca: Optimize Resource Allocation for Data-parallel Jobs using Deep LearningZhiyao Hu, Dongsheng Li, Dongxiang Zhang, Yixin ChenINFOCOM 2020 · 3 citations
Related papers
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
- Unlocking unallocated cloud capacity for long, uninterruptible workloadsAnup Agarwal, Shadi A. Noghabi, Íñigo Goiri, Srinivasan Seshan et al.NSDI 2023 · 9 citations
- A Case for Task Sampling based Learning for Cluster Job SchedulingAkshay Jajoo, Y. Charlie Hu, Xiaojun Lin, Nan DengNSDI 2022 · 35 citations
- Generating Complex, Realistic Cloud Workloads using Recurrent Neural NetworksShane Bergsma, Timothy Zeyl, Arik Senderovich, J. Christopher BeckSOSP 2021 · 21 citations
- Caerus: NIMBLE Task Scheduling for Serverless AnalyticsHong Zhang, Yupeng Tang, Anurag Khandelwal, Jingrong Chen et al.NSDI 2021 · 75 citations
