PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
Jahid Hasan, Stanley Jiang, Tejendra Singh, Sainyam Galhotra, Romila Pradhan, Divesh Srivastava
Abstract
Data is a critical component of modern decision-making systems; system malfunctions (e.g., performance degradation and module failure) can often be traced back to a mismatch between the properties of the data and the assumptions of the system modules that process the data. For example, with the increasing use of open-source libraries to develop data science pipelines, common causes of system malfunctions include inappropriately configured data processing libraries for data cleaning tasks such as entity resolution or missing value imputation. Our objective is to resolve malfunctioning pipelines and improve their utility; we introduce PipeLens, a framework that leverages successful and unsuccessful runs of past pipelines for fixing pipeline malfunctions. PipeLens uses an acyclic graph representation of the pipeline and performs causal reasoning through interventions: when a system malfunctions with a given dataset, PipeLens modifies the pipeline (by changing its structure or the parameters of its modules) and observes the impact of this intervention on system behavior. To focus on useful interventions, we learn a proxy function that approximates the pipeline's utility over a dataset and guides the search for the best intervention. Unlike traditional observational analysis that reports correlations between system parameters and their behavior, we provide causally verified root causes and suggest pipeline modifications that rectify malfunctions. Empirical evaluation on four data science tasks over four real-world datasets demonstrates that PipeLens consistently outperforms baselines in terms of interventions performed to repair malfunctions while maintaining practical running times.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d9e62f6-63ba-4432-b6a5-508bd5dd0b70Builds on10
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel et al.VLDB 2021 · 69 citations
- The Art and Practice of Data Science Pipelines: A Comprehensive Study of Data Science Pipelines In Theory, In-The-Small, and In-The-LargeSumon Biswas, Mohammad Wardat, Hridesh RajanICSE 2022 · 64 citations
- Probabilistic Delta debuggingGuancheng Wang, Ruobing Shen, Junjie Chen, Yingfei Xiong et al.FSE 2021 · 56 citations
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
Related papers
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 9 citations
- DataPrism: Exposing Disconnect between Data and SystemsSainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire et al.SIGMOD 2022 · 9 citations
- Data Debugging with Shapley Importance over Machine Learning PipelinesBojan Karlas, David Dao, Matteo Interlandi, Sebastian Schelter et al.ICLR 2024 · 11 citations
- Can Machine Learning Pipelines Be Better Configured?Yibo Wang, Ying Wang, Tingwei Zhang, Yue Yu et al.FSE 2023 · 4 citations
- Stress-Testing ML Pipelines with Adversarial Data CorruptionJiongli Zhu, Geyang Xu, Felipe Lorenzi, Boris Glavic et al.VLDB 2025 · 2 citations
