Data Glitches Discovery using Influence-based Model Explanations
Nikolaos Myrtakis, Ioannis Tsamardinos, Vassilis Christophides
Abstract
We address the problem of detecting data glitches in ML training sets, specifically mislabeled and anomalous samples. Detection of data glitches provides insights into the quality of the data sampling. Their repair may improve the reliability and the performance of the model. The proposed methodology is based on exploiting influence functions that estimate how much the loss of the model (or a given sample) is affected when a sample is removed from the training set. We introduce three novel signals for detecting, characterizing, and repairing data glitches in a training set based on sample influences. Influence-based signals form an explainable-by-design data glitch detection framework, producing intuitively explainable signals of the actual predictive model built. In contrast, specialized algorithms that are agnostic to the target ML model (e.g., anomaly detectors) replicate the work of fitting the data distribution and may detect glitches that are inconsistent with the decision boundary of the predictive model. Computational experiments on tabular and image data modalities demonstrate that the proposed signals outperform, in some cases up to a factor of 6, all existing influence-based signals, and generalize across different datasets and ML models. In addition, they often outperform specialized glitch detectors (e.g., mislabeled and anomaly detectors) and provide accurate label repairs for mislabeled samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51177013-93cf-4e32-b9a0-f70d2d282cd3Cited by top-tier papers2
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- EDDI: Explaining Data Drift Using InfluenceNikolaos Myrtakis, Andrea Castellani, Ioannis Tsamardinos, Vassilis ChristophidesICDE 2026
Builds on14
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Scaling Up Influence FunctionsAndrea Schioppa, Polina Zablotskaia, David Vilar, Artem SokolovAAAI 2022 · 149 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
Related papers
- Repairing Neural Networks by Leaving the Right Past BehindRyutaro Tanno, Melanie F. Pradier, Aditya V. Nori, Yingzhen LiNeurIPS 2022 · 41 citations
- Resolving Training Biases via Influence-based Data RelabelingShuming Kong, Yanyan Shen, Linpeng HuangICLR 2022 · 71 citations
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang et al.VLDB 2024 · 13 citations
- Debugging and Explaining Metric Learning Approaches: An Influence Function Based PerspectiveRuofan Liu, Yun Lin, Xianglin Yang, Jin Song DongNeurIPS 2022 · 4 citations
- Complaint-driven Training Data Debugging for Query 2.0Weiyuan Wu, Lampros Flokas, Eugene Wu, Jiannan WangSIGMOD 2020 · 36 citations
