Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines
Sijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas, Divesh Srivastava
Abstract
Despite the increasing success of Machine Learning (ML) techniques in real-world applications, their maintenance over time remains challenging. In particular, the prediction accuracy of deployed ML models can suffer due to significant changes between training and serving data over time, known as data drift. Traditional data drift solutions primarily focus on detecting drift, and then retraining the ML models, but do not discern whether the detected drift is harmful to model performance. In this paper, we observe that not all data drifts lead to degradation in prediction accuracy. We then introduce a novel approach for identifying portions of data distributions in serving data where drift can be potentially harmful to model performance, which we term Data Distributions with Low Accuracy (DDLA). Our approach, using decision trees, precisely pinpoints low-accuracy zones within ML models, especially Blackbox models. By focusing on these DDLAs, we effectively assess the impact of data drift on model performance and make informed decisions in the ML pipeline. In contrast to existing data drift techniques, we advocate for model retraining only in cases of harmful drifts that detrimentally affect model performance. Through extensive experimental evaluations on various datasets and models, our findings demonstrate that our approach significantly improves cost-efficiency over baselines, while achieving comparable accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1797034f-dc85-46af-b801-028b4d7deed5Cited by top-tier papers3
- Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on InstagramAnna Ricarda Luther, Hendrik Heuer, Sebastian Haunss, Stephanie Geise et al.CHI 2026 · 1 citation
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance DriftHarvineet Singh, Fan Xia, Alexej Gossmann, Andrew Chuang et al.ICML 2025
- EDDI: Explaining Data Drift Using InfluenceNikolaos Myrtakis, Andrea Castellani, Ioannis Tsamardinos, Vassilis ChristophidesICDE 2026
Builds on5
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel et al.VLDB 2021 · 69 citations
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 36 citations
- Comparing Distributions by Measuring Differences that Affect Decision MakingShengjia Zhao, Abhishek Sinha, Yutong He, Aidan Perreault et al.ICLR 2022 · 28 citations
- A Learning Based Hypothesis Test for Harmful Covariate ShiftTom Ginsberg, Zhongyuan Liang, Rahul G. KrishnanICLR 2023 · 3 citations
Related papers
- Detecting Interpretable Subgroup DriftsFlavio Giobergia, Eliana Pastor, Luca de Alfaro, Elena BaralisKDD 2025
- Reliably detecting model failures in deployment without labelsViet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya et al.NeurIPS 2025 · 3 citations
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- When to retrain a machine learning modelFlorence Regol, Leo Schwinn, Kyle Sprague, Mark Coates et al.ICML 2025
- Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsGwanjong Park, Osama Khan, Dongho Ha, Myeongjae Jeon et al.EuroSys 2026 · 1 citation
