Data Debugging with Shapley Importance over Machine Learning Pipelines
Bojan Karlas, David Dao, Matteo Interlandi, Sebastian Schelter, Wentao Wu, Ce Zhang
摘要
When a machine learning (ML) model exhibits poor quality (e.g., poor accuracy or fairness), the problem can often be traced back to errors in the training data. Being able to discover the data examples that are the most likely culprits is a fundamental concern that has received a lot of attention recently. One prominent way to measure "data importance" with respect to model quality is the Shapley value. Unfortunately, existing methods only focus on the ML model in isolation, without considering the broader ML pipeline for data preparation and feature extraction, which appears in the majority of real-world ML code. This presents a major limitation to applying existing methods in practical settings. In this paper, we propose Datascope, a method for efficiently computing Shapley-based data importance over ML pipelines. We introduce several approximations that lead to dramatic improvements in terms of computational speed. Finally, our experimental evaluation demonstrates that our methods are capable of data error discovery that is as effective as existing Monte Carlo baselines, and in some cases even outperform them. We release our code as an open-source data debugging library available at github.com/easeml/datascope.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Shapley Value Estimation based on Differential MatrixJunyuan Pang, Jian Pei, Haocheng Xia, Xiang Li 等SIGMOD 2025 · 被引用 2 次
- Reliable and Private Utility Signaling for Data MarketsLi Peng, Jiayao Zhang, Yihang Wu, Weiran Liu 等SIGMOD 2026 · 被引用 1 次
- Local Shapley: Model-Induced Locality and Optimal Reuse in Data ValuationXuan Yang, Hsi-Wen Chen, Ming-Syan Chen, Jian PeiVLDB 2026 · 被引用 1 次
- Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang 等SIGMOD 2026
它引用的顶会 Paper6
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 被引用 36 次
相关 Paper
- Capturing and querying fine-grained provenance of preprocessing pipelines in data scienceAdriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo TorloneVLDB 2021 · 被引用 39 次
- Scalability vs. Utility: Do We Have To Sacrifice One for the Other in Data Importance Quantification?Ruoxi Jia, Fan Wu, Xuehui Sun, Jiacen Xu 等CVPR 2021
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 被引用 54 次
- CS-Shapley: Class-wise Shapley Values for Data Valuation in ClassificationStephanie Schoch, Haifeng Xu, Yangfeng JiNeurIPS 2022 · 被引用 56 次
- WeShap: Weak Supervision Source Evaluation with Shapley ValuesNaiqing Guan, Nick KoudasVLDB 2025
