Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines
Sijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas, Divesh Srivastava
摘要
Despite the increasing success of Machine Learning (ML) techniques in real-world applications, their maintenance over time remains challenging. In particular, the prediction accuracy of deployed ML models can suffer due to significant changes between training and serving data over time, known as data drift. Traditional data drift solutions primarily focus on detecting drift, and then retraining the ML models, but do not discern whether the detected drift is harmful to model performance. In this paper, we observe that not all data drifts lead to degradation in prediction accuracy. We then introduce a novel approach for identifying portions of data distributions in serving data where drift can be potentially harmful to model performance, which we term Data Distributions with Low Accuracy (DDLA). Our approach, using decision trees, precisely pinpoints low-accuracy zones within ML models, especially Blackbox models. By focusing on these DDLAs, we effectively assess the impact of data drift on model performance and make informed decisions in the ML pipeline. In contrast to existing data drift techniques, we advocate for model retraining only in cases of harmful drifts that detrimentally affect model performance. Through extensive experimental evaluations on various datasets and models, our findings demonstrate that our approach significantly improves cost-efficiency over baselines, while achieving comparable accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on InstagramAnna Ricarda Luther, Hendrik Heuer, Sebastian Haunss, Stephanie Geise 等CHI 2026 · 被引用 1 次
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance DriftHarvineet Singh, Fan Xia, Alexej Gossmann, Andrew Chuang 等ICML 2025
- EDDI: Explaining Data Drift Using InfluenceNikolaos Myrtakis, Andrea Castellani, Ioannis Tsamardinos, Vassilis ChristophidesICDE 2026
它引用的顶会 Paper5
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang 等ICML 2020 · 被引用 213 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 被引用 36 次
- Comparing Distributions by Measuring Differences that Affect Decision MakingShengjia Zhao, Abhishek Sinha, Yutong He, Aidan Perreault 等ICLR 2022 · 被引用 28 次
- A Learning Based Hypothesis Test for Harmful Covariate ShiftTom Ginsberg, Zhongyuan Liang, Rahul G. KrishnanICLR 2023 · 被引用 3 次
相关 Paper
- Detecting Interpretable Subgroup DriftsFlavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena BaralisKDD 2025
- Reliably detecting model failures in deployment without labelsViet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya 等NeurIPS 2025 · 被引用 3 次
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- When to retrain a machine learning modelFlorence Regol, Leo Schwinn, Kyle Sprague, Mark Coates 等ICML 2025
- Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsGwanjong Park, Osama Khan, Dongho Ha, Myeongjae Jeon 等EuroSys 2026 · 被引用 1 次
