DataPrism: Exposing Disconnect between Data and Systems
Sainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire, Alexandra Meliou, Divesh Srivastava
Abstract
As data is a central component of many modern systems, the cause of a system malfunction may reside in the data, and, specifically, particular properties of data. E.g., a health-monitoring system that is designed under the assumption that weight is reported in lbs will malfunction when encountering weight reported in kilograms. Like software debugging, which aims to find bugs in the source code or runtime conditions, our goal is to debug data to identify potential sources of disconnect between the assumptions about some data and systems that operate on that data. We propose DataPrism, a framework to identify data properties (profiles) that are the root causes of performance degradation or failure of a data-driven system. Such identification is necessary to repair data and resolve the disconnect between data and systems. Our technique is based on causal reasoning through interventions: when a system malfunctions for a dataset, DataPrism alters the data profiles and observes changes in the system's behavior due to the alteration. Unlike statistical observational analysis that reports mere correlations, DataPrism reports causally verified root causes -- in terms of data profiles -- of the system malfunction. We empirically evaluate DataPrism on seven real-world and several synthetic data-driven systems that fail on certain datasets due to a diverse set of reasons. In all cases, DataPrism identifies the root causes precisely while requiring orders of magnitude fewer interventions than prior techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1954afe-db8f-4d9d-b20e-1cd182b89471Cited by top-tier papers3
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 9 citations
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 3 citations
- PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science PipelinesJahid Hasan, Stanley Jiang, Tejendra Singh, Sainyam Galhotra et al.VLDB 2026
Builds on7
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Causal Relational LearningBabak Salimi, Harsh Parikh, Moe Kayali, Lise Getoor et al.SIGMOD 2020 · 38 citations
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 36 citations
- SCODED: Statistical Constraint Oriented Data Error DetectionJing Nathan Yan, Oliver Schulte, Mohan Zhang, Jiannan Wang et al.SIGMOD 2020 · 32 citations
- Causality-Guided Adaptive Interventional DebuggingAnna Fariha, Suman Nath, Alexandra MeliouSIGMOD 2020 · 19 citations
Related papers
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 9 citations
- Trading Personalization for Accuracy: Data Debugging in Collaborative FilteringLong Chen, Yuan Yao, Feng Xu, Miao Xu et al.NeurIPS 2020 · 8 citations
- Stress-Testing Causal Claims via Cardinality RepairsYarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig et al.SIGMOD 2026 · 1 citation
- What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingChenyang Yang, Yining Hong, Grace A. Lewis, Tongshuang Wu et al.ASE 2024 · 2 citations
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
