Tracking the risk of a deployed model and detecting harmful distribution shifts
Aleksandr Podkopaev, Aaditya Ramdas
Abstract
When deployed in the real world, machine learning models inevitably encounter changes in the data distribution, and certain-but not all-distribution shifts could result in significant performance degradation. In practice, it may make sense to ignore benign shifts, under which the performance of a deployed model does not degrade substantially, making interventions by a human expert (or model retraining) unnecessary. While several works have developed tests for distribution shifts, these typically either use non-sequential methods, or detect arbitrary shifts (benign or harmful), or both. We argue that a sensible method for firing off a warning has to both (a) detect harmful shifts while ignoring benign ones, and (b) allow continuous monitoring of model performance without increasing the false alarm rate. In this work, we design simple sequential tools for testing if the difference between source (training) and target (test) distributions leads to a significant increase in a risk function of interest, like accuracy or calibration. Recent advances in constructing time-uniform confidence sequences allow efficient aggregation of statistical evidence accumulated during the tracking process. The designed framework is applicable in settings where (some) true labels are revealed after the prediction is performed, or when batches of labels become available in a delayed fashion. We demonstrate the efficacy of the proposed framework through an extensive empirical study on a collection of simulated and real datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65569ba1-2b28-4b51-b26f-e70070beec95Cited by top-tier papers15
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 37 citations
- Sequential Covariate Shift Detection Using Classifier Two-Sample TestsSooyong Jang, Sangdon Park, Insup Lee, Osbert BastaniICML 2022 · 24 citations
- Sequential Changepoint Detection via Backward Confidence SequencesShubhanshu Shekhar, Aaditya RamdasICML 2023 · 16 citations
- Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesSijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas et al.VLDB 2024 · 13 citations
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.NeurIPS 2024 · 13 citations
Builds on3
- Classification with Valid and Adaptive CoverageYaniv Romano, Matteo Sesia, Emmanuel J. CandèsNeurIPS 2020 · 586 citations
- Distribution-free binary classification: prediction sets, confidence intervals and calibrationChirag Gupta, Aleksandr Podkopaev, Aaditya RamdasNeurIPS 2020 · 105 citations
- Top-label calibration and multiclass-to-binary reductionsChirag Gupta, Aaditya RamdasICLR 2022 · 51 citations
Related papers
- Monitoring Risks in Test-Time AdaptationMona Schirmer, Metod Jazbec, Christian Andersson Naesseth, Eric T. NalisnickNeurIPS 2025 · 10 citations
- WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal MartingalesDrew Prinster, Xing Han, Anqi Liu, Suchi SariaICML 2025
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment SettingsAngéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas PapernotICML 2025
- Auditing Fairness by BettingBen Chugg, Santiago Cortes-Gomez, Bryan Wilder, Aaditya RamdasNeurIPS 2023 · 29 citations
- A Learning Based Hypothesis Test for Harmful Covariate ShiftTom Ginsberg, Zhongyuan Liang, Rahul G. KrishnanICLR 2023 · 3 citations
