Lune

NeurIPS2025顶会

Reliably detecting model failures in deployment without labels

Viet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya, Amol Verma, Fahad Razak, Rahul G. Krishnan

2025年份
3被引次数

摘要

The distribution of data changes over time; models operating in dynamic environments need retraining. But knowing when to retrain, without access to labels, is an open challenge since some, but not all shifts degrade model performance. This paper formalizes and addresses the problem of post-deployment deterioration (PDD) monitoring. We propose D3M, a practical and efficient monitoring algorithm based on the disagreement of predictive models, achieving low false positive rates under non-deteriorating shifts and provides sample complexity bounds for high true positive rates under deteriorating shifts. Empirical results on both standard benchmark and a real-world large-scale internal medicine dataset demonstrate the effectiveness of the framework and highlight its viability as an alert mechanism for high-stakes machine learning pipelines.

2 Background and Algorithm

Assume a base model f is supervisedly trained to classify inputs x ∈ X into finite discrete classes Y = 1, . . . , C from training examples D n = x i , y i i=1:n where tuples (x i , y i ) i=1:n ∼ P n for all i ∈ [n]. For a joint distribution P over X × Y, let P x denote its marginal distribution over X . We are interested in designing a mechanism such that upon seeing a collection of unlabeled inputs x ′ i i=1:m sampled from a deployment distribution Q x , the mechanism flags model deterioration if Q x ̸ = P x and f underperforms on Q x without being able to observe labels for Q x . On the other hand, if Q x ̸ = P x while f remains performant on Q x , the mechanism should resist flagging. Achieving so ensures that our mechanism only flags deployment-time changes that are truly deteriorating.

How can we monitor ML models for deployment deterioration due to distribution shift without indiscriminately flagging any changes in the data?

We require a computable quantity ϕ, independent of labels, whose value statistically differs indistribution (ID) and out-of-distribution (OOD) if and only if model deterioration occurs. Monitoring, then, regresses to recording baseline values for ϕ evaluated on ID held-out validation samples. Then, upon collecting unsupervised deployment samples from an unknown distribution, the monitoring mechanism computes φ and compares it to the recorded baseline values, and finally outputs a verdict.

Leveraging insights from [16,24,19], the framework of model disagreement possesses this property under certain assumptions about the underlying distribution change (see Appendix A). We say that two models h 1 and h 2 disagree on an input x ∈ X if h 1 (x) ̸ = h 2 (x). In particular, models exhibit greater predictive disagreement on unsupervised samples that lead to model deterioration, compared to in-distribution (ID) samples. This is observed through the increased entropy-based discrepancy between classification heads in [16], or maximum disagreement between models in the same ensemble in [24] and [19], as signal for detecting deployment deterioration. Maximizing classification head discrepancy for OOD detection [16] is efficient for monitoring at deployment time, requiring only one forward pass to compute a verdict. However, this trades off classification accuracy as this training procedure alters the original trained decision boundaries. On the other hand, computing model disagreement between ensembles [19] requires finetuning a potentially large network to collect ID and deployment-time disagreement statistics ϕ. In addition, ID training data is required at deployment, further limiting the scalability of such framework.

To avoid needing the original training set at deployment, we replace finetuning with a Bayesian approach that models a posterior predictive distribution (PPD) over logits. This yields a distribution over decision boundaries that remains faithful to ID behavior. By comparing samples from the PPD to the mean prediction, we approximate maximum disagreement without retraining or access to training data. As the PPD is usually intractable, we instead model it with a variational distribution, easily optimizable using standard methods.

Sampling disagreement statistics ϕ in this way yields a reference distribution Φ of ID maximum disagreement rates. At deployment, we compute the same statistic φ and flag model deterioration when φ exceeds a high quantile of Φ. This enables unsupervised, training-free monitoring. Our method-Disagreement-Driven Deterioration Monitoring (D3M)-follows three key steps: Train, Calibrate, and Deploy.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext e6c89043-e00a-4a90-8290-8b27f448be5e

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖