PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
Jahid Hasan, Stanley Jiang, Tejendra Singh, Sainyam Galhotra, Romila Pradhan, Divesh Srivastava
摘要
Data is a critical component of modern decision-making systems; system malfunctions (e.g., performance degradation and module failure) can often be traced back to a mismatch between the properties of the data and the assumptions of the system modules that process the data. For example, with the increasing use of open-source libraries to develop data science pipelines, common causes of system malfunctions include inappropriately configured data processing libraries for data cleaning tasks such as entity resolution or missing value imputation. Our objective is to resolve malfunctioning pipelines and improve their utility; we introduce PipeLens, a framework that leverages successful and unsuccessful runs of past pipelines for fixing pipeline malfunctions. PipeLens uses an acyclic graph representation of the pipeline and performs causal reasoning through interventions: when a system malfunctions with a given dataset, PipeLens modifies the pipeline (by changing its structure or the parameters of its modules) and observes the impact of this intervention on system behavior. To focus on useful interventions, we learn a proxy function that approximates the pipeline's utility over a dataset and guides the search for the best intervention. Unlike traditional observational analysis that reports correlations between system parameters and their behavior, we provide causally verified root causes and suggest pipeline modifications that rectify malfunctions. Empirical evaluation on four data science tasks over four real-world datasets demonstrates that PipeLens consistently outperforms baselines in terms of interventions performed to repair malfunctions while maintaining practical running times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- The Art and Practice of Data Science Pipelines: A Comprehensive Study of Data Science Pipelines In Theory, In-The-Small, and In-The-LargeSumon Biswas, Mohammad Wardat, Hridesh RajanICSE 2022 · 被引用 64 次
- Probabilistic Delta debuggingGuancheng Wang, Ruobing Shen, Junjie Chen, Yingfei Xiong 等FSE 2021 · 被引用 56 次
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 被引用 24 次
相关 Paper
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 被引用 9 次
- DataPrism: Exposing Disconnect between Data and SystemsSainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire 等SIGMOD 2022 · 被引用 9 次
- Data Debugging with Shapley Importance over Machine Learning PipelinesBojan Karlas, David Dao, Matteo Interlandi, Sebastian Schelter 等ICLR 2024 · 被引用 11 次
- Can Machine Learning Pipelines Be Better Configured?Yibo Wang, Ying Wang, Tingwei Zhang, Yue Yu 等FSE 2023 · 被引用 4 次
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic 等VLDB 2025 · 被引用 2 次
