DataPrism: Exposing Disconnect between Data and Systems
Sainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire, Alexandra Meliou, Divesh Srivastava
摘要
As data is a central component of many modern systems, the cause of a system malfunction may reside in the data, and, specifically, particular properties of data. E.g., a health-monitoring system that is designed under the assumption that weight is reported in lbs will malfunction when encountering weight reported in kilograms. Like software debugging, which aims to find bugs in the source code or runtime conditions, our goal is to debug data to identify potential sources of disconnect between the assumptions about some data and systems that operate on that data. We propose DataPrism, a framework to identify data properties (profiles) that are the root causes of performance degradation or failure of a data-driven system. Such identification is necessary to repair data and resolve the disconnect between data and systems. Our technique is based on causal reasoning through interventions: when a system malfunctions for a dataset, DataPrism alters the data profiles and observes changes in the system's behavior due to the alteration. Unlike statistical observational analysis that reports mere correlations, DataPrism reports causally verified root causes -- in terms of data profiles -- of the system malfunction. We empirically evaluate DataPrism on seven real-world and several synthetic data-driven systems that fail on certain datasets due to a diverse set of reasons. In all cases, DataPrism identifies the root causes precisely while requiring orders of magnitude fewer interventions than prior techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 被引用 9 次
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 被引用 3 次
- PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science PipelinesJahid Hasan, Stanley Jiang, Tejendra Singh, Sainyam Galhotra 等VLDB 2026
它引用的顶会 Paper7
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- Causal Relational LearningBabak Salimi, Harsh Parikh, Moe Kayali, Lise Getoor 等SIGMOD 2020 · 被引用 38 次
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 被引用 36 次
- SCODED: Statistical Constraint Oriented Data Error DetectionJing Nathan Yan, Oliver Schulte, Mohan Zhang, Jiannan Wang 等SIGMOD 2020 · 被引用 32 次
- Causality-Guided Adaptive Interventional DebuggingAnna Fariha, Suman Nath, Alexandra MeliouSIGMOD 2020 · 被引用 19 次
相关 Paper
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 被引用 9 次
- Trading Personalization for Accuracy: Data Debugging in Collaborative FilteringLong Chen, Yuan Yao, Feng Xu, Miao Xu 等NeurIPS 2020 · 被引用 8 次
- Stress-Testing Causal Claims via Cardinality RepairsYarden Gabbay, Haoquan Guan, Shaull Almagor, El Kindi Rezig 等SIGMOD 2026 · 被引用 1 次
- What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingChenyang Yang, Yining Hong, Grace A. Lewis, Tongshuang Wu 等ASE 2024 · 被引用 2 次
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
