EDDI: Explaining Data Drift Using Influence
Nikolaos Myrtakis, Andrea Castellani, Ioannis Tsamardinos, Vassilis Christophides
摘要
Data drift poses a significant challenge to the reliability of ML models in real-world applications, as the data distribution may change between training and inference phases. Existing drift monitoring systems have several limitations: (i) they detect either model drift or feature drift but not both; (ii) they rely on a plethora of detection methods, each making its own assumptions regarding the data and the ML models; (iii) they do not explain a drift in relation to the underlying probability distribution. To address these limitations, we propose EDDI, a novel influence-based drift detection and explanation framework that leverages the direct influence of samples on the decision boundary of the deployed predictive model. EDDI represents the first unified framework that utilizes Influence Functions to detect both model and feature drift. At the same time, EDDI provides a novel kind of explanation by revealing the probabilistic source of a drift. The key idea is that the influence distributions between drifted and non-drifted samples differ. Moreover, different drift types exhibit unique distributional influence characteristics, enabling their attribution. Based on a type-specific drift simulation, it is demonstrated that EDDI achieves statistically significant improvements on detection performance up to 15% over baselines across four drift types on 42 time-series classification datasets and reveals the drift type with up to 0.8 AUC. Moreover, we evaluated EDDI on two datasets with naturally occurring drift.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai 等ICML 2024 · 被引用 442 次
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang 等USENIX Security 2017 · 被引用 325 次
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi 等USENIX Security 2021 · 被引用 241 次
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
相关 Paper
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- Towards Non-Parametric Drift Detection via Dynamic Adapting Window Independence Drift Detection (DAWIDD)Fabian Hinder, André Artelt, Barbara HammerICML 2020 · 被引用 28 次
- Context-Aware Drift DetectionOliver Cobb, Arnaud Van LooverenICML 2022 · 被引用 22 次
- Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesSijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas 等VLDB 2024 · 被引用 13 次
- Early Concept Drift Detection via Prediction UncertaintyPengqian Lu, Jie Lu, Anjin Liu, Guangquan ZhangAAAI 2025 · 被引用 12 次
