A Learning Based Hypothesis Test for Harmful Covariate Shift
Tom Ginsberg, Zhongyuan Liang, Rahul G. Krishnan
Abstract
The ability to quickly and accurately identify covariate shift at test time is a critical and often overlooked component of safe machine learning systems deployed in high-risk domains. While methods exist for detecting when predictions should not be made on out-of-distribution test examples, identifying distributional level differences between training and test time can help determine when a model should be removed from the deployment setting and retrained. In this work, we define harmful covariate shift (HCS) as a change in distribution that may weaken the generalization of a predictive model. To detect HCS, we use the discordance between an ensemble of classifiers trained to agree on training data and disagree on test data. We derive a loss function for training this ensemble and show that the disagreement rate and entropy represent powerful discriminative statistics for HCS. Empirically, we demonstrate the ability of our method to detect harmful covariate shift with statistical certainty on a variety of high-dimensional datasets. Across numerous domains and modalities, we show state-of-the-art performance compared to existing methods, particularly when the number of observed test samples is small 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4426f4fb-44ab-47da-9b7f-cafb81b0c274Cited by top-tier papers10
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 37 citations
- A Geometric Explanation of the Likelihood OOD Detection ParadoxHamidreza Kamkari, Brendan Leigh Ross, Jesse C. Cresswell, Anthony L. Caterini et al.ICML 2024 · 20 citations
- Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesSijie Dong, Qitong Wang, Soror Sahri, Themis Palpanas et al.VLDB 2024 · 13 citations
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Adaptive Uncertainty Estimation via High-Dimensional Testing on Latent RepresentationsTsai Hor Chan, Kin Wai Lau, Jiajun Shen, Guosheng Yin et al.NeurIPS 2023 · 3 citations
Builds on8
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Detecting Out-of-Distribution Examples with Gram MatricesChandramouli Shama Sastry, Sageev OoreICML 2020 · 275 citations
- Learning Deep Kernels for Non-Parametric Two-Sample TestsFeng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang et al.ICML 2020 · 213 citations
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
- Understanding Failures in Out-of-Distribution Detection with Deep Generative ModelsLily H. Zhang, Mark Goldstein, Rajesh RanganathICML 2021 · 129 citations
Related papers
- Sequential Covariate Shift Detection Using Classifier Two-Sample TestsSooyong Jang, Sangdon Park, Insup Lee, Osbert BastaniICML 2022 · 24 citations
- Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment SettingsAngéline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas PapernotICML 2025
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 36 citations
- Unified Out-Of-Distribution Detection: A Model-Specific PerspectiveReza Averly, Wei-Lun ChaoICCV 2023 · 18 citations
- (Almost) Provable Error Bounds Under Distribution Shift via Disagreement DiscrepancyElan Rosenfeld, Saurabh GargNeurIPS 2023 · 18 citations
