LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems
Arnab Phani, Benjamin Rath, Matthias Boehm
摘要
Machine learning (ML) and data science workflows are inherently exploratory. Data scientists pose hypotheses, integrate the necessary data, and run ML pipelines of data cleaning, feature engineering, model selection and hyper-parameter tuning. The repetitive nature of these workflows, and their hierarchical composition from building blocks exhibits high computational redundancy. Existing work addresses this redundancy with coarse-grained lineage tracing and reuse for ML pipelines. This approach allows using existing ML systems, but views entire algorithms as black boxes, and thus, fails to eliminate fine-grained redundancy and to handle internal non-determinism. In this paper, we introduce LIMA, a practical framework for efficient, fine-grained lineage tracing and reuse inside ML systems. Lineage tracing of individual operations creates new challenges and opportunities. We address the large size of lineage traces with multi-level lineage tracing and reuse, as well as lineage deduplication for loops and functions; exploit full and partial reuse opportunities across the program hierarchy; and integrate this framework with task parallelism and operator fusion. The resulting framework performs fine-grained lineage tracing with low overhead, provides versioning and reproducibility, and is able to eliminate fine-grained redundancy. Our experiments on a variety of ML pipelines show performance improvements up to 12.4x.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 被引用 13 次
- ElasticNotebook: Enabling Live Migration for Computational NotebooksZhaoheng Li, Pranav Gor, Rahul Prabhu, Hui Yu 等VLDB 2024 · 被引用 12 次
- Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training DatasetsSupun Nakandala, Arun KumarSIGMOD 2022 · 被引用 6 次
- CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML PipelinesSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2025 · 被引用 5 次
它引用的顶会 Paper6
- Fine-Grained Lineage for Safer Notebook InteractionsStephen Macke, Aditya G. Parameswaran, Hongpu Gong, Doris Jung Lin Lee 等VLDB 2021 · 被引用 46 次
- Vamsa: Automated Provenance Tracking in Data Science ScriptsMohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas, Subru Krishnan 等KDD 2020 · 被引用 41 次
- Optimizing DNN Computation Graph using Graph SubstitutionsJingzhi Fang, Yanyan Shen, Yue Wang, Lei ChenVLDB 2020 · 被引用 29 次
- Optimizing Machine Learning Workloads in Collaborative EnvironmentsBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl 等SIGMOD 2020 · 被引用 22 次
- SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear AlgebraYisu Remy Wang, Shana Hutchison, Dan Suciu, Bill Howe 等VLDB 2020
相关 Paper
- Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGsXiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar 等SIGMOD 2025 · 被引用 1 次
- Capturing and querying fine-grained provenance of preprocessing pipelines in data scienceAdriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo TorloneVLDB 2021 · 被引用 39 次
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló 等ICDE 2024 · 被引用 1 次
- Reproducible ContainersOmar S. Navarro Leija, Kelly Shiptoski, Ryan G. Scott, Baojun Wang 等ASPLOS 2020 · 被引用 21 次
- Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning PipelinesStefan Grafberger, Paul Groth, Sebastian SchelterSIGMOD 2023 · 被引用 18 次
