Estimating and Explaining Model Performance When Both Covariates and Labels Shift
Lingjiao Chen, Matei Zaharia, James Y. Zou
Abstract
Deployed machine learning (ML) models often encounter new user data that differs from their training data. Therefore, estimating how well a given model might perform on the new data is an important step toward reliable ML applications. This is very challenging, however, as the data distribution can change in flexible ways, and we may not have any labels on the new data, which is often the case in monitoring settings. In this paper, we propose a new distribution shift model, Sparse Joint Shift (SJS), which considers the joint shift of both labels and a few features. This unifies and generalizes several existing shift models including label shift and sparse covariate shift, where only marginal feature or label distribution shifts are considered. We describe mathematical conditions under which SJS is identifiable. We further propose SEES, an algorithmic framework to characterize the distribution shift under SJS and to estimate a model's performance on new data without any labels. We conduct extensive experiments on several real-world datasets with various ML models. Across different datasets and distribution shifts, SEES achieves significant (up to an order of magnitude) shift estimation error improvements over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcb7b507-cec0-408c-ada7-d4e610d990c0Cited by top-tier papers6
- Adapting to Continuous Covariate Shift via Online Density Ratio EstimationYu-Jie Zhang, Zhen-Yu Zhang, Peng Zhao, Masashi SugiyamaNeurIPS 2023 · 25 citations
- Energy-based Automated Model EvaluationRu Peng, Heming Zou, Haobo Wang, Yawen Zeng et al.ICLR 2024 · 18 citations
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag et al.NeurIPS 2025 · 9 citations
- Fed-ADE: Adaptive Learning Rate for Federated Post-adaptation under Distribution ShiftHeewon Park, Mugon Joe, Miru Kim, Kyungjin Im et al.CVPR 2026 · 2 citations
- Explaining Concept Shift with Interpretable Feature AttributionRuiqi Lyu, Alistair Turcan, Bryan WilderICML 2026
Builds on15
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann et al.NeurIPS 2020 · 688 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
Related papers
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Estimating Model Performance Under Covariate Shift Without LabelsJakub Bialek, Juhani Kivimäki, Wojtek Kuberski, Nikolaos PerrakisNeurIPS 2025 · 10 citations
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 51 citations
- Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and DetectionHaoyue Bai, Gregory Canal, Xuefeng Du, Jeongyeol Kwon et al.ICML 2023 · 67 citations
