Estimating Model Performance Under Covariate Shift Without Labels
Jakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos Perrakis
摘要
After deployment, machine learning models often experience performance degradation due to shifts in data distribution. It is challenging to assess post-deployment performance accurately when labels are missing or delayed. Existing proxy methods, such as data drift detection, fail to measure the effects of these shifts adequately. To address this, we introduce a new method for evaluating binary classification models on unlabeled tabular data that accurately estimates model performance under covariate shift and call it Probabilistic Adaptive Performance Estimation (PAPE). It can be applied to any performance metric defined with elements of the confusion matrix. Crucially, PAPE operates independently of the original model, relying only on its predictions and probability estimates, and does not need any assumptions about the nature of covariate shift, learning directly from data instead. We tested PAPE using over 900 dataset-model combinations from the US census data, assessing its performance against several benchmarks through various metrics. Our findings show that PAPE outperforms other methodologies, making it a superior choice for estimating the performance of binary classification models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur 等ICLR 2022 · 被引用 160 次
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell 等ICCV 2021 · 被引用 141 次
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 被引用 134 次
相关 Paper
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Characterizing Out-of-Distribution Error via Optimal TransportYuzhe Lu, Yilong Qin, Runtian Zhai, Andrew Shen 等NeurIPS 2023 · 被引用 22 次
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment SettingsAngéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas PapernotICML 2025
- Estimating Generalization under Distribution Shifts via Domain-Invariant RepresentationsChing-Yao Chuang, Antonio Torralba, Stefanie JegelkaICML 2020 · 被引用 72 次
- Monitoring Risks in Test-Time AdaptationMona Schirmer, Metod Jazbec, Christian Andersson Naesseth, Eric T. NalisnickNeurIPS 2025 · 被引用 10 次
