Reliably detecting model failures in deployment without labels
Viet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya, Amol Verma, Fahad Razak, Rahul G. Krishnan
Abstract
The distribution of data changes over time; models operating in dynamic environments need retraining. But knowing when to retrain, without access to labels, is an open challenge since some, but not all shifts degrade model performance. This paper formalizes and addresses the problem of post-deployment deterioration (PDD) monitoring. We propose D3M, a practical and efficient monitoring algorithm based on the disagreement of predictive models, achieving low false positive rates under non-deteriorating shifts and provides sample complexity bounds for high true positive rates under deteriorating shifts. Empirical results on both standard benchmark and a real-world large-scale internal medicine dataset demonstrate the effectiveness of the framework and highlight its viability as an alert mechanism for high-stakes machine learning pipelines.
2 Background and Algorithm
Assume a base model f is supervisedly trained to classify inputs x ∈ X into finite discrete classes Y = 1, . . . , C from training examples D n = x i , y i i=1:n where tuples (x i , y i ) i=1:n ∼ P n for all i ∈ [n]. For a joint distribution P over X × Y, let P x denote its marginal distribution over X . We are interested in designing a mechanism such that upon seeing a collection of unlabeled inputs x ′ i i=1:m sampled from a deployment distribution Q x , the mechanism flags model deterioration if Q x ̸ = P x and f underperforms on Q x without being able to observe labels for Q x . On the other hand, if Q x ̸ = P x while f remains performant on Q x , the mechanism should resist flagging. Achieving so ensures that our mechanism only flags deployment-time changes that are truly deteriorating.
How can we monitor ML models for deployment deterioration due to distribution shift without indiscriminately flagging any changes in the data?
We require a computable quantity ϕ, independent of labels, whose value statistically differs indistribution (ID) and out-of-distribution (OOD) if and only if model deterioration occurs. Monitoring, then, regresses to recording baseline values for ϕ evaluated on ID held-out validation samples. Then, upon collecting unsupervised deployment samples from an unknown distribution, the monitoring mechanism computes φ and compares it to the recorded baseline values, and finally outputs a verdict.
Leveraging insights from [16,24,19], the framework of model disagreement possesses this property under certain assumptions about the underlying distribution change (see Appendix A). We say that two models h 1 and h 2 disagree on an input x ∈ X if h 1 (x) ̸ = h 2 (x). In particular, models exhibit greater predictive disagreement on unsupervised samples that lead to model deterioration, compared to in-distribution (ID) samples. This is observed through the increased entropy-based discrepancy between classification heads in [16], or maximum disagreement between models in the same ensemble in [24] and [19], as signal for detecting deployment deterioration. Maximizing classification head discrepancy for OOD detection [16] is efficient for monitoring at deployment time, requiring only one forward pass to compute a verdict. However, this trades off classification accuracy as this training procedure alters the original trained decision boundaries. On the other hand, computing model disagreement between ensembles [19] requires finetuning a potentially large network to collect ID and deployment-time disagreement statistics ϕ. In addition, ID training data is required at deployment, further limiting the scalability of such framework.
To avoid needing the original training set at deployment, we replace finetuning with a Bayesian approach that models a posterior predictive distribution (PPD) over logits. This yields a distribution over decision boundaries that remains faithful to ID behavior. By comparing samples from the PPD to the mean prediction, we approximate maximum disagreement without retraining or access to training data. As the PPD is usually intractable, we instead model it with a variational distribution, easily optimizable using standard methods.
Sampling disagreement statistics ϕ in this way yields a reference distribution Φ of ID maximum disagreement rates. At deployment, we compute the same statistic φ and flag model deterioration when φ exceeds a high quantile of Φ. This enables unsupervised, training-free monitoring. Our method-Disagreement-Driven Deterioration Monitoring (D3M)-follows three key steps: Train, Calibrate, and Deploy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6c89043-e00a-4a90-8290-8b27f448be5eBuilds on17
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 193 citations
Related papers
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment SettingsAngéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas PapernotICML 2025
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Estimating Model Performance Under Covariate Shift Without LabelsJakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos PerrakisNeurIPS 2025 · 10 citations
- Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesSijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas et al.VLDB 2024 · 13 citations
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 36 citations
