Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications
Hani Al-Sayeh, Bunjamin Memishi, Muhammad Attahir Jibril, Marcus Paradies, Kai-Uwe Sattler
摘要
Distributed in-memory processing frameworks accelerate iterative workloads by caching suitable datasets in memory rather than recomputing them in each iteration. Selecting appropriate datasets to cache as well as allocating a suitable cluster configuration for caching these datasets play a crucial role in achieving optimal performance. In practice, both are tedious, time-consuming tasks and are often neglected by end users, who are typically not aware of workload semantics, sizes of intermediate data, and cluster specification.
To address these problems, we present Juggler, an end-to-end framework, which autonomously selects appropriate datasets for caching and recommends a correspondingly suitable cluster configuration to end users, with the aim of achieving optimal execution time and cost. We evaluate Juggler on various iterative, real-world, machine learning applications. Compared with our baseline, Juggler reduces execution time to 25.1 % and cost to 58.1 %, on average, as a result of selecting suitable datasets for caching. It recommends optimal cluster configuration in 50 % of cases and near-to-optimal configuration in the remaining cases. Moreover, Juggler achieves an average performance prediction accuracy of 90 %.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- Black or White? How to Develop an AutoTuner for Memory-based AnalyticsMayuresh Kunjir, Shivnath BabuSIGMOD 2020 · 被引用 65 次
- OPTIMUSCLOUD: Heterogeneous Configuration Optimization for Distributed Databases in the CloudAshraf Mahgoub, Alexander Medoff, Rakesh Kumar, Subrata Mitra 等USENIX ATC 2020 · 被引用 63 次
- On the Use of ML for Blackbox System Performance PredictionSilvery Fu, Saurabh Gupta, Radhika Mittal, Sylvia RatnasamyNSDI 2021 · 被引用 43 次
- Detecting cache-related bugs in Spark applicationsHui Li, Dong Wang, Tianze Huang, Yu Gao 等ISSTA 2020 · 被引用 8 次
相关 Paper
- Blaze: Holistic Caching for Iterative Data ProcessingWon Wook Song, Jeongyoon Eo, Taegeon Um, Myeongjae Jeon 等EuroSys 2024 · 被引用 1 次
- Partitioner Selection with EASE to Optimize Distributed Graph ProcessingNikolai Merkel, Ruben Mayer, Tawkir Ahmed Fakir, Hans-Arno JacobsenICDE 2023 · 被引用 3 次
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha 等ICDE 2021 · 被引用 15 次
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 被引用 91 次
- LSched: A Workload-Aware Learned Query Scheduler for Analytical Database SystemsIbrahim Sabek, Tenzin Samten Ukyab, Tim KraskaSIGMOD 2022 · 被引用 25 次
