Optimizing Machine Learning Workloads in Collaborative Environments
Behrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl, Volker Markl
Abstract
Effective collaboration among data scientists results in high-quality and efficient machine learning (ML) workloads. In a collaborative environment, such as Kaggle or Google Colabratory, users typically re-execute or modify published scripts to recreate or improve the result. This introduces many redundant data processing and model training operations. Reusing the data generated by the redundant operations leads to the more efficient execution of future workloads. However, existing collaborative environments lack a data management component for storing and reusing the result of previously executed operations. In this paper, we present a system to optimize the execution of ML workloads in collaborative environments by reusing previously performed operations and their results. We utilize a so-called Experiment Graph (EG) to store the artifacts, i.e., raw and intermediate data or ML models, as vertices and operations of ML workloads as edges. In theory, the size of EG can become unnecessarily large, while the storage budget might be limited. At the same time, for some artifacts, the overall storage and retrieval cost might outweigh the recomputation cost. To address this issue, we propose two algorithms for materializing artifacts based on their likelihood of future reuse. Given the materialized artifacts inside EG, we devise a linear-time reuse algorithm to find the optimal execution plan for incoming ML workloads. Our reuse algorithm only incurs a negligible overhead and scales for the high number of incoming ML workloads in collaborative environments. Our experiments show that we improve the run-time by one order of magnitude for repeated execution of the workloads and 50% for the execution of modified workloads in collaborative environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b22a103d-40f9-4277-a2ed-9118ef6f0e57Cited by top-tier papers11
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak et al.VLDB 2021 · 35 citations
- ModelKeeper: Accelerating DNN Training via Automated Training WarmupFan Lai, Yinwei Dai, Harsha V. Madhyastha, Mosharaf ChowdhuryNSDI 2023 · 31 citations
- MLCask: Efficient Management of Component Evolution in Collaborative Data Analytics PipelinesZhaojing Luo, Sai Ho Yeung, Meihui Zhang, Kaiping Zheng et al.ICDE 2021 · 31 citations
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 30 citations
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 13 citations
Related papers
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló et al.ICDE 2024 · 1 citation
- Materialization and Reuse Optimizations for Production Data Science PipelinesBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Zoi Kaoudi, Tilmann Rabl et al.SIGMOD 2022 · 12 citations
- KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceMossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier et al.ICDE 2024 · 7 citations
- Towards Scalable Online Machine Learning Collaborations with OpenMLJoaquin VanschorenVLDB 2021 · 1 citation
- Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGsXiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar et al.SIGMOD 2025 · 1 citation
