Phoebe: A Learning-based Checkpoint Optimizer
Yiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das, Hiren Patel, Malay Bag, Hitesh Sharma, Alekh Jindal
摘要
Easy-to-use programming interfaces paired with cloud-scale processing engines have enabled big data system users to author arbitrarily complex analytical jobs over massive volumes of data. However, as the complexity and scale of analytical jobs increase, they encounter a number of unforeseen problems, hotspots with large intermediate data on temporary storage, longer job recovery time after failures, and worse query optimizer estimates being examples of issues that we are facing at Microsoft. To address these issues, we propose Phoebe, an efficient learning-based checkpoint optimizer. Given a set of constraints and an objective function at compile-time, Phoebe is able to determine the decomposition of job plans, and the optimal set of checkpoints to preserve their outputs to durable global storage. Phoebe consists of three machine learning predictors and one optimization module. For each stage of a job, Phoebe makes accurate predictions for: (1) the execution time, (2) the output size, and (3) the start/end time taking into account the inter-stage dependencies. Using these predictions, we formulate checkpoint optimization as an integer programming problem and propose a scalable heuristic algorithm that meets the latency requirement of the production environment. We demonstrate the effectiveness of Phoebe in production workloads, and show that we can free the temporary storage on hotspots by more than 70% and restart failed jobs 68% faster on average with minimum performance impact. Phoebe also illustrates that adding multiple sets of checkpoints is not cost-efficient, which dramatically reduces the complexity of the optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha 等VLDB 2022 · 被引用 14 次
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 被引用 5 次
它引用的顶会 Paper4
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- ML-based Cross-Platform Query OptimizationZoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Bertty Contreras-Rojas, Rodrigo Pardo-Meza 等ICDE 2020 · 被引用 27 次
- Incorporating Super-Operators in Big-Data Query OptimizersJyoti Leeka, Kaushik RajanVLDB 2020 · 被引用 17 次
相关 Paper
- Practical Parameterized Query Optimization via Efficient Plan Reuse and List-wise RankingHai Lan, Yang Yu, Zhifeng Bao, Zi Huang 等SIGMOD 2026
- Towards Optimizing Storage Costs on the CloudKoyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh 等ICDE 2023 · 被引用 8 次
- ELENA: AN Explainability-Aided Online Query Optimization FrameworkYuan Dong, Yuanyuan Yao, Yangyang Wu, Lu Chen 等ICDE 2026
- A Resource-Aware Deep Cost Model for Big Data Query ProcessingYan Li, Liwei Wang, Sheng Wang, Yuan Sun 等ICDE 2022 · 被引用 13 次
- Eraser: Eliminating Performance Regression on Learned Query OptimizerLianggui Weng, Rong Zhu, Di Wu, Bolin Ding 等VLDB 2024 · 被引用 19 次
