Reliably detecting model failures in deployment without labels
Viet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya, Amol Verma, Fahad Razak, Rahul G. Krishnan
摘要
The distribution of data changes over time; models operating in dynamic environments need retraining. But knowing when to retrain, without access to labels, is an open challenge since some, but not all shifts degrade model performance. This paper formalizes and addresses the problem of post-deployment deterioration (PDD) monitoring. We propose D3M, a practical and efficient monitoring algorithm based on the disagreement of predictive models, achieving low false positive rates under non-deteriorating shifts and provides sample complexity bounds for high true positive rates under deteriorating shifts. Empirical results on both standard benchmark and a real-world large-scale internal medicine dataset demonstrate the effectiveness of the framework and highlight its viability as an alert mechanism for high-stakes machine learning pipelines.
2 Background and Algorithm
Assume a base model f is supervisedly trained to classify inputs x ∈ X into finite discrete classes Y = 1, . . . , C from training examples D n = x i , y i i=1:n where tuples (x i , y i ) i=1:n ∼ P n for all i ∈ [n]. For a joint distribution P over X × Y, let P x denote its marginal distribution over X . We are interested in designing a mechanism such that upon seeing a collection of unlabeled inputs x ′ i i=1:m sampled from a deployment distribution Q x , the mechanism flags model deterioration if Q x ̸ = P x and f underperforms on Q x without being able to observe labels for Q x . On the other hand, if Q x ̸ = P x while f remains performant on Q x , the mechanism should resist flagging. Achieving so ensures that our mechanism only flags deployment-time changes that are truly deteriorating.
How can we monitor ML models for deployment deterioration due to distribution shift without indiscriminately flagging any changes in the data?
We require a computable quantity ϕ, independent of labels, whose value statistically differs indistribution (ID) and out-of-distribution (OOD) if and only if model deterioration occurs. Monitoring, then, regresses to recording baseline values for ϕ evaluated on ID held-out validation samples. Then, upon collecting unsupervised deployment samples from an unknown distribution, the monitoring mechanism computes φ and compares it to the recorded baseline values, and finally outputs a verdict.
Leveraging insights from [16,24,19], the framework of model disagreement possesses this property under certain assumptions about the underlying distribution change (see Appendix A). We say that two models h 1 and h 2 disagree on an input x ∈ X if h 1 (x) ̸ = h 2 (x). In particular, models exhibit greater predictive disagreement on unsupervised samples that lead to model deterioration, compared to in-distribution (ID) samples. This is observed through the increased entropy-based discrepancy between classification heads in [16], or maximum disagreement between models in the same ensemble in [24] and [19], as signal for detecting deployment deterioration. Maximizing classification head discrepancy for OOD detection [16] is efficient for monitoring at deployment time, requiring only one forward pass to compute a verdict. However, this trades off classification accuracy as this training procedure alters the original trained decision boundaries. On the other hand, computing model disagreement between ensembles [19] requires finetuning a potentially large network to collect ID and deployment-time disagreement statistics ϕ. In addition, ID training data is required at deployment, further limiting the scalability of such framework.
To avoid needing the original training set at deployment, we replace finetuning with a Bayesian approach that models a posterior predictive distribution (PPD) over logits. This yields a distribution over decision boundaries that remains faithful to ID behavior. By comparing samples from the PPD to the mean prediction, we approximate maximum disagreement without retraining or access to training data. As the PPD is usually intractable, we instead model it with a variational distribution, easily optimizable using standard methods.
Sampling disagreement statistics ϕ in this way yields a reference distribution Φ of ID maximum disagreement rates. At deployment, we compute the same statistic φ and flag model deterioration when φ exceeds a high quantile of Φ. This enables unsupervised, training-free monitoring. Our method-Disagreement-Driven Deterioration Monitoring (D3M)-follows three key steps: Train, Calibrate, and Deploy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang 等ICML 2020 · 被引用 213 次
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 被引用 193 次
相关 Paper
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment SettingsAngéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas PapernotICML 2025
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Estimating Model Performance Under Covariate Shift Without LabelsJakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos PerrakisNeurIPS 2025 · 被引用 10 次
- Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesSijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas 等VLDB 2024 · 被引用 13 次
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 被引用 36 次
