Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
Angéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas Papernot
摘要
Deploying machine learning models in safetycritical domains poses a key challenge: ensuring reliable model performance on downstream user data without access to ground truth labels for direct validation. We propose the suitability filter, a novel framework designed to detect performance deterioration by utilizing suitability signals-model output features that are sensitive to covariate shifts and indicative of potential prediction errors. The suitability filter evaluates whether classifier accuracy on unlabeled user data shows significant degradation compared to the accuracy measured on the labeled test dataset. Specifically, it ensures that this degradation does not exceed a pre-specified margin, which represents the maximum acceptable drop in accuracy. To achieve reliable performance evaluation, we aggregate suitability signals for both test and user data and compare these empirical distributions using statistical hypothesis testing, thus providing insights into decision uncertainty. Our modular method adapts to various models and domains. Empirical evaluations across different classification tasks demonstrate that the suitability filter reliably detects performance deviations due to covariate shift. This enables proactive mitigation of potential failures in high-stakes applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 被引用 986 次
- A Fine-Grained Analysis on Distribution ShiftOlivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi 等ICLR 2022 · 被引用 258 次
- Self-Adaptive Training: beyond Empirical Risk MinimizationLang Huang, Chao Zhang, Hongyang ZhangNeurIPS 2020 · 被引用 256 次
相关 Paper
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Estimating Model Performance Under Covariate Shift Without LabelsJakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos PerrakisNeurIPS 2025 · 被引用 10 次
- A Learning Based Hypothesis Test for Harmful Covariate ShiftTom Ginsberg, Zhongyuan Liang, Rahul G. KrishnanICLR 2023 · 被引用 3 次
- Reliably detecting model failures in deployment without labelsViet Nguyen, Changjian Shui, Vijay Giri, Siddharth Arya 等NeurIPS 2025 · 被引用 3 次
- Unified Out-Of-Distribution Detection: A Model-Specific PerspectiveReza Averly, Wei-Lun ChaoICCV 2023 · 被引用 18 次
