HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine Learning
Antonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló, Sergi Nadal
Abstract
We present HYPPO, a novel system to optimize pipelines encountered in exploratory machine learning. HYPPO exploits alternative computational paths of artifacts from past executions to derive better execution plans while reusing materialized artifacts. Adding alternative computations introduces new challenges for exploratory machine learning regarding workload representation, system architecture, and optimal execution plan generation. To this end, we present a novel workload representation based on directed hypergraphs, and we formulate the problem of discovering the optimal execution plan as a search problem over directed hypergraphs and that of selecting artifacts to materialize as an optimization problem. A thorough experimental evaluation shows that HYPPO results in plans that are typically one order (up to two orders) of magnitude faster and cheaper than the non-optimized pipeline and considerably (up to one order of magnitude) faster and cheaper than plans generated by the state of the art when materializing artifacts is possible. Lastly, our evaluation reveals that HYPPO reduces the cost by 3–4× even when materialization cannot be exploited.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9be0681e-9ed1-40a9-bb1f-9c2bc41d373fCited by top-tier papers1
Ask how each one uses itBuilds on17
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang et al.NeurIPS 2022 · 157 citations
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan et al.VLDB 2023 · 127 citations
Related papers
- Optimizing Machine Learning Workloads in Collaborative EnvironmentsBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl et al.SIGMOD 2020 · 22 citations
- Materialization and Reuse Optimizations for Production Data Science PipelinesBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Zoi Kaoudi, Tilmann Rabl et al.SIGMOD 2022 · 12 citations
- Hippo: Sharing Computations in Hyper-Parameter OptimizationAhnjae Shin, Joo Seong Jeong, Do Yoon Kim, Soyoung Jung et al.VLDB 2022 · 6 citations
- Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGsXiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar et al.SIGMOD 2025 · 1 citation
- EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized ViewsZhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, Umakishore RamachandranSIGMOD 2022 · 26 citations
