LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems
Arnab Phani, Benjamin Rath, Matthias Boehm
Abstract
Machine learning (ML) and data science workflows are inherently exploratory. Data scientists pose hypotheses, integrate the necessary data, and run ML pipelines of data cleaning, feature engineering, model selection and hyper-parameter tuning. The repetitive nature of these workflows, and their hierarchical composition from building blocks exhibits high computational redundancy. Existing work addresses this redundancy with coarse-grained lineage tracing and reuse for ML pipelines. This approach allows using existing ML systems, but views entire algorithms as black boxes, and thus, fails to eliminate fine-grained redundancy and to handle internal non-determinism. In this paper, we introduce LIMA, a practical framework for efficient, fine-grained lineage tracing and reuse inside ML systems. Lineage tracing of individual operations creates new challenges and opportunities. We address the large size of lineage traces with multi-level lineage tracing and reuse, as well as lineage deduplication for loops and functions; exploit full and partial reuse opportunities across the program hierarchy; and integrate this framework with task parallelism and operator fusion. The resulting framework performs fine-grained lineage tracing with low overhead, provides versioning and reproducibility, and is able to eliminate fine-grained redundancy. Our experiments on a variety of ML pipelines show performance improvements up to 12.4x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 13 citations
- ElasticNotebook: Enabling Live Migration for Computational NotebooksZhaoheng Li, Pranav Gor, Rahul Prabhu, Hui Yu et al.VLDB 2024 · 12 citations
- Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training DatasetsSupun Nakandala, Arun KumarSIGMOD 2022 · 6 citations
- CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML PipelinesSaeed Fathollahzadeh, Essam Mansour, Matthias BoehmVLDB 2025 · 5 citations
Builds on6
- Fine-Grained Lineage for Safer Notebook InteractionsStephen Macke, Aditya G. Parameswaran, Hongpu Gong, Doris Jung Lin Lee et al.VLDB 2021 · 46 citations
- Vamsa: Automated Provenance Tracking in Data Science ScriptsMohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas, Subru Krishnan et al.KDD 2020 · 41 citations
- Optimizing DNN Computation Graph using Graph SubstitutionsJingzhi Fang, Yanyan Shen, Yue Wang, Lei ChenVLDB 2020 · 29 citations
- Optimizing Machine Learning Workloads in Collaborative EnvironmentsBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl et al.SIGMOD 2020 · 22 citations
- SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear AlgebraYisu Remy Wang, Shana Hutchison, Dan Suciu, Bill Howe et al.VLDB 2020
Related papers
- Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGsXiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar et al.SIGMOD 2025 · 1 citation
- Capturing and querying fine-grained provenance of preprocessing pipelines in data scienceAdriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo TorloneVLDB 2021 · 39 citations
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló et al.ICDE 2024 · 1 citation
- Reproducible ContainersOmar S. Navarro Leija, Kelly Shiptoski, Ryan G. Scott, Baojun Wang et al.ASPLOS 2020 · 21 citations
- Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning PipelinesStefan Grafberger, Paul Groth, Sebastian SchelterSIGMOD 2023 · 18 citations
