Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGs
Xiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar, Sadeem Alsudais, Chen Li
Abstract
Data analytics tasks are often formulated as data workflows represented as directed acyclic graphs (DAGs) of operators. The recent trend of adopting machine learning (ML) techniques in workflows results in increasingly complicated DAGs with many operators and edges. Compared to the operator-at-a-time execution paradigm, pipelined execution has benefits of reducing the materialization cost of intermediate results and allowing operators to produce results early, which are critical in iterative analysis on large data volumes. Correctly scheduling a workflow DAG for pipelined execution is non-trivial due to the richer semantics of operators and the increasing complexity of DAGs. Several existing data systems adopt simple heuristics to solve the problem without considering costs such as materialization sizes. In this paper, we systematically study the problem of scheduling a workflow DAG for pipelined execution, and develop a novel cost-based optimizer called Pasta for generating a high-quality schedule. The Pasta optimizer is not only general and applicable to a wide variety of cost functions, but also capable of utilizing properties inherent in a broad class of cost functions to improve its performance significantly. We conducted a thorough evaluation of developed techniques on real-world workflows and show the efficiency and efficacy of these solutions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 89bfc4d6-af36-409d-9192-b5f56fc7a719Related papers
- Searching for Machine Learning Pipelines Using a Context-Free GrammarRadu Marinescu, Akihiro Kishimoto, Parikshit Ram, Ambrish Rawat et al.AAAI 2021 · 18 citations
- Data Flow Lifecycles for Optimizing Workflow CoordinationHyungro Lee, Luanzheng Guo, Meng Tang, Jesun Firoz et al.SC 2023 · 8 citations
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló et al.ICDE 2024 · 1 citation
- Materialization and Reuse Optimizations for Production Data Science PipelinesBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Zoi Kaoudi, Tilmann Rabl et al.SIGMOD 2022 · 12 citations
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 30 citations
