MLCask: Efficient Management of Component Evolution in Collaborative Data Analytics Pipelines
Zhaojing Luo, Sai Ho Yeung, Meihui Zhang, Kaiping Zheng, Lei Zhu, Gang Chen, Feiyi Fan, Qian Lin, Kee Yuan Ngiam, Beng Chin Ooi
Abstract
With the ever-increasing adoption of machine learning for data analytics, maintaining a machine learning pipeline is becoming more complex as both the datasets and trained models evolve with time. In a collaborative environment, the changes and updates due to pipeline evolution often cause cumbersome coordination and maintenance work, raising the costs and making it hard to use. Existing solutions, unfortunately, do not address the version evolution problem, especially in a collaborative environment where non-linear version control semantics are necessary to isolate operations made by different user roles. The lack of version control semantics also incurs unnecessary storage consumption and lowers efficiency due to data duplication and repeated data pre-processing, which are avoidable.
In this paper, we identify two main challenges that arise during the deployment of machine learning pipelines, and address them with the design of versioning for an end-to-end analytics system MLCask. The system supports multiple user roles with the ability to perform Git-like branching and merging operations in the context of the machine learning pipelines. We define and accelerate the metric-driven merge operation by pruning the pipeline search tree using reusable history records and pipeline compatibility information. Further, we design and implement the prioritized pipeline search, which gives preference to the pipelines that probably yield better performance. The effectiveness of MLCask is evaluated through an extensive study over several real-world deployment cases. The performance evaluation shows that the proposed merge operation is up to 7.8x faster and saves up to 11.9x storage space than the baseline method that does not utilize history records.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fd6bcdc-c1cc-45da-a490-a7ca881af627Cited by top-tier papers10
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang et al.CVPR 2022 · 115 citations
- Falcon: A Privacy-Preserving and Interpretable Vertical Federated Learning SystemYuncheng Wu, Naili Xing, Gang Chen, Tien Tuan Anh Dinh et al.VLDB 2023 · 47 citations
- Serverless Data Science - Are We There Yet? A Case Study of Model ServingYuncheng Wu, Tien Tuan Anh Dinh, Guoyu Hu, Meihui Zhang et al.SIGMOD 2022 · 27 citations
- AlphaEvolve: A Learning Framework to Discover Novel Alphas in Quantitative InvestmentCan Cui, Wei Wang, Meihui Zhang, Gang Chen et al.SIGMOD 2021 · 25 citations
- FEAST: A Communication-efficient Federated Feature Selection Framework for Relational DataRui Fu, Yuncheng Wu, Quanqing Xu, Meihui ZhangSIGMOD 2023 · 17 citations
Builds on3
- TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes ApplicationsKaiping Zheng, Shaofeng Cai, Horng Ruey Chua, Wei Wang et al.SIGMOD 2020 · 22 citations
- Optimizing Machine Learning Workloads in Collaborative EnvironmentsBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl et al.SIGMOD 2020 · 22 citations
- Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource ConstraintsShaofeng Cai, Gang Chen, Beng Chin Ooi, Jinyang GaoVLDB 2020 · 21 citations
Related papers
- MGit: A Model Versioning and Management SystemWei Hao, Daniel Mendoza, Rafael Mendes, Deepak Narayanan et al.ICML 2024 · 1 citation
- Enabling Secure and Efficient Data Analytics Pipeline Evolution with Trusted Execution EnvironmentHaotian Gao, Cong Yue, Tien Tuan Anh Dinh, Zhiyong Huang et al.VLDB 2023 · 6 citations
- Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning PipelinesStefan Grafberger, Paul Groth, Sebastian SchelterSIGMOD 2023 · 18 citations
- Git-Theta: A Git Extension for Collaborative Development of Machine Learning ModelsNikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas et al.ICML 2023 · 15 citations
- Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and ProcessNadia Nahar, Shurui Zhou, Grace A. Lewis, Christian KästnerICSE 2022 · 122 citations
