Phoebe: A Learning-based Checkpoint Optimizer
Yiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das, Hiren Patel, Malay Bag, Hitesh Sharma, Alekh Jindal
Abstract
Easy-to-use programming interfaces paired with cloud-scale processing engines have enabled big data system users to author arbitrarily complex analytical jobs over massive volumes of data. However, as the complexity and scale of analytical jobs increase, they encounter a number of unforeseen problems, hotspots with large intermediate data on temporary storage, longer job recovery time after failures, and worse query optimizer estimates being examples of issues that we are facing at Microsoft. To address these issues, we propose Phoebe, an efficient learning-based checkpoint optimizer. Given a set of constraints and an objective function at compile-time, Phoebe is able to determine the decomposition of job plans, and the optimal set of checkpoints to preserve their outputs to durable global storage. Phoebe consists of three machine learning predictors and one optimization module. For each stage of a job, Phoebe makes accurate predictions for: (1) the execution time, (2) the output size, and (3) the start/end time taking into account the inter-stage dependencies. Using these predictions, we formulate checkpoint optimization as an integer programming problem and propose a scalable heuristic algorithm that meets the latency requirement of the production environment. We demonstrate the effectiveness of Phoebe in production workloads, and show that we can free the temporary storage on hotspots by more than 70% and restart failed jobs 68% faster on average with minimum performance impact. Phoebe also illustrates that adding multiple sets of checkpoints is not cost-efficient, which dramatically reduces the complexity of the optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 428a8803-bc6b-496d-9032-39bfdd262bbfCited by top-tier papers2
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha et al.VLDB 2022 · 14 citations
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 5 citations
Builds on4
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang et al.VLDB 2021 · 138 citations
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
- ML-based Cross-Platform Query OptimizationZoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Bertty Contreras-Rojas, Rodrigo Pardo-Meza et al.ICDE 2020 · 27 citations
- Incorporating Super-Operators in Big-Data Query OptimizersJyoti Leeka, Kaushik RajanVLDB 2020 · 17 citations
Related papers
- Practical Parameterized Query Optimization via Efficient Plan Reuse and List-wise RankingHai Lan, Yang Yu, Zhifeng Bao, Zi Huang et al.SIGMOD 2026
- Towards Optimizing Storage Costs on the CloudKoyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh et al.ICDE 2023 · 8 citations
- ELENA: AN Explainability-Aided Online Query Optimization FrameworkYuan Dong, Yuanyuan Yao, Yangyang Wu, Lu Chen et al.ICDE 2026
- A Resource-Aware Deep Cost Model for Big Data Query ProcessingYan Li, Liwei Wang, Sheng Wang, Yuan Sun et al.ICDE 2022 · 13 citations
- Eraser: Eliminating Performance Regression on Learned Query OptimizerLianggui Weng, Rong Zhu, Di Wu, Bolin Ding et al.VLDB 2024 · 19 citations
