Capturing and querying fine-grained provenance of preprocessing pipelines in data science
Adriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo Torlone
摘要
Data processing pipelines that are designed to clean, transform and alter data in preparation for learning predictive models, have an impact on those models' accuracy and performance, as well on other properties, such as model fairness. It is therefore important to provide developers with the means to gain an in-depth understanding of how the pipeline steps affect the data, from the raw input to training sets ready to be used for learning. While other efforts track creation and changes of pipelines of relational operators, in this work we analyze the typical operations of data preparation within a machine learning process, and provide infrastructure for generating very granular provenance records from it, at the level of individual elements within a dataset. Our contributions include: (i) the formal definition of a core set of preprocessing operators, and the definition of provenance patterns for each of them, and (ii) a prototype implementation of an application-level provenance capture library that works alongside Python. We report on provenance processing and storage overhead and scalability experiments, carried out over both real ML benchmark pipelines and over TCP-DI, and show how the resulting provenance can be used to answer a suite of provenance benchmark queries that underpin some of the developers' debugging questions, as expressed on the Data Science Stack Exchange.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Observability for Production Machine Learning Pipelines [Vision]Shreya Shankar, Aditya G. ParameswaranVLDB 2022 · 被引用 21 次
- Modyn: Data-Centric Machine Learning Pipeline OrchestrationMaximilian Böther, Ties Robroek, Viktor Gsteiger, Robin Holzinger 等SIGMOD 2025 · 被引用 6 次
- Toward Temporal Attribution Analytics in Dataflows [Vision Paper]Chrysanthi Kosyfaki, Ruiyuan Zhang, Nikos Mamoulis, Xiaofang ZhouVLDB 2026
它引用的顶会 Paper3
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke 等VLDB 2020 · 被引用 109 次
- PrIU: A Provenance-Based Approach for Incrementally Updating Regression ModelsYinjun Wu, Val Tannen, Susan B. DavidsonSIGMOD 2020 · 被引用 26 次
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 被引用 9 次
相关 Paper
- Data Debugging with Shapley Importance over Machine Learning PipelinesBojan Karlas, David Dao, Matteo Interlandi, Sebastian Schelter 等ICLR 2024 · 被引用 11 次
- Vamsa: Automated Provenance Tracking in Data Science ScriptsMohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas, Subru Krishnan 等KDD 2020 · 被引用 41 次
- Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipelineSumon Biswas, Hridesh RajanFSE 2021 · 被引用 101 次
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 被引用 30 次
- Experimental Analysis of Multi-Step Pipelines for Fair Classifications - More than the Sum of Their Parts?Nico Lässig, Melanie HerschelICDE 2025
