EDDI: Explaining Data Drift Using Influence
Nikolaos Myrtakis, Andrea Castellani, Ioannis Tsamardinos, Vassilis Christophides
Abstract
Data drift poses a significant challenge to the reliability of ML models in real-world applications, as the data distribution may change between training and inference phases. Existing drift monitoring systems have several limitations: (i) they detect either model drift or feature drift but not both; (ii) they rely on a plethora of detection methods, each making its own assumptions regarding the data and the ML models; (iii) they do not explain a drift in relation to the underlying probability distribution. To address these limitations, we propose EDDI, a novel influence-based drift detection and explanation framework that leverages the direct influence of samples on the decision boundary of the deployed predictive model. EDDI represents the first unified framework that utilizes Influence Functions to detect both model and feature drift. At the same time, EDDI provides a novel kind of explanation by revealing the probabilistic source of a drift. The key idea is that the influence distributions between drifted and non-drifted samples differ. Moreover, different drift types exhibit unique distributional influence characteristics, enabling their attribution. Based on a type-specific drift simulation, it is demonstrated that EDDI achieves statistically significant improvements on detection performance up to 15% over baselines across four drift types on 42 time-series classification datasets and reveals the drift type with up to 0.8 AUC. Moreover, we evaluated EDDI on two datasets with naturally occurring drift.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c570db9-fb52-4cb2-94d1-10c85261d80eBuilds on9
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai et al.ICML 2024 · 442 citations
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang et al.USENIX Security 2017 · 325 citations
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi et al.USENIX Security 2021 · 241 citations
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
Related papers
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- Towards Non-Parametric Drift Detection via Dynamic Adapting Window Independence Drift Detection (DAWIDD)Fabian Hinder, André Artelt, Barbara HammerICML 2020 · 28 citations
- Context-Aware Drift DetectionOliver Cobb, Arnaud Van LooverenICML 2022 · 22 citations
- Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesSijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas et al.VLDB 2024 · 13 citations
- Early Concept Drift Detection via Prediction UncertaintyPengqian Lu, Jie Lu, Anjin Liu, Guangquan ZhangAAAI 2025 · 12 citations
