A Unified View of Label Shift Estimation
Saurabh Garg, Yifan Wu, Sivaraman Balakrishnan, Zachary C. Lipton
Abstract
Label shift describes the setting where although the label distribution might change between the source and target domains, the class-conditional probabilities (of data given a label) do not. There are two dominant approaches for estimating the label marginal. BBSE, a moment-matching approach based on confusion matrices, is provably consistent and provides interpretable error bounds. However, a maximum likelihood estimation approach, which we call MLLS, dominates empirically. In this paper, we present a unified view of the two methods and the first theoretical characterization of the likelihood-based estimator. Our contributions include (i) conditions for consistency of MLLS, which include calibration of the classifier and a confusion matrix invertibility condition that BBSE also requires; (ii) a unified view of the methods, casting the confusion matrix as roughly equivalent to MLLS for a particular choice of calibration method; and (iii) a decomposition of MLLS's finite-sample error into terms reflecting the impacts of miscalibration and estimation error. Our analysis attributes BBSE's statistical inefficiency to a loss of information due to coarse calibration. We support our findings with experiments on both synthetic data and the MNIST and CIFAR10 image recognition datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfc6ba28-8e7f-4f9b-bd37-42d523b59478Cited by top-tier papers56
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep LearningJannik Kossen, Neil Band, Clare Lyle, Aidan N. Gomez et al.NeurIPS 2021 · 180 citations
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
- Maximum Likelihood with Bias-Corrected Calibration is Hard-To-Beat at Label Shift AdaptationAmr Alexandari, Anshul Kundaje, Avanti ShrikumarICML 2020 · 123 citations
- Distribution-free binary classification: prediction sets, confidence intervals and calibrationChirag Gupta, Aleksandr Podkopaev, Aaditya RamdasNeurIPS 2020 · 105 citations
Related papers
- LaSCal: Label-Shift Calibration without target labelsTeodora Popordanoska, Gorjan Radevski, Tinne Tuytelaars, Matthew B. BlaschkoNeurIPS 2024 · 12 citations
- Expectation Consistency Loss: Rethink Confidence Calibration under Covariate ShiftJinzong Dong, Zhaohui Jiang, Bo YangICML 2026
- ELSA: Efficient Label Shift Adaptation through the Lens of Semiparametric ModelsQinglong Tian, Xin Zhang, Jiwei ZhaoICML 2023 · 12 citations
- Label Shift Correction via Bidirectional Marginal Distribution MatchingRuidong Fan, Xiao Ouyang, Hong Tao, Chenping HouKDD 2024 · 1 citation
- Graph-Smoothed Bayesian Black-Box Shift Estimator and Its Information GeometryMasanari KimuraNeurIPS 2025 · 1 citation
