Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training Ensembles
Jiefeng Chen, Frederick Liu, Besim Avci, Xi Wu, Yingyu Liang, Somesh Jha
Abstract
When a deep learning model is deployed in the wild, it can encounter test data drawn from distributions different from the training data distribution and suffer drop in performance. For safe deployment, it is essential to estimate the accuracy of the pre-trained model on the test data. However, the labels for the test inputs are usually not immediately available in practice, and obtaining them can be expensive. This observation leads to two challenging tasks: (1) unsupervised accuracy estimation, which aims to estimate the accuracy of a pre-trained classifier on a set of unlabeled test inputs; (2) error detection, which aims to identify mis-classified test inputs. In this paper, we propose a principled and practically effective framework that simultaneously addresses the two tasks. The proposed framework iteratively learns an ensemble of models to identify mis-classified data points and performs self-training to improve the ensemble with the identified points. Theoretical analysis demonstrates that our framework enjoys provable guarantees for both accuracy estimation and error detection under mild conditions readily satisfied by practical deep learning models. Along with the framework, we proposed and experimented with two instantiations and achieved state-of-the-art results on 59 tasks. For example, on iWildCam, one instantiation reduces the estimation error for unsupervised accuracy estimation by at least 70% and improves the F1 score for error detection by at least 4.7% compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c07f845-2c9e-494c-bd3f-bb91314df67aCited by top-tier papers25
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 134 citations
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 120 citations
- Estimating and Explaining Model Performance When Both Covariates and Labels ShiftLingjiao Chen, Matei Zaharia, James Y. ZouNeurIPS 2022 · 34 citations
- On the Strong Correlation Between Model Invariance and GeneralizationWeijian Deng, Stephen Gould, Liang ZhengNeurIPS 2022 · 28 citations
Builds on7
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 134 citations
- Mandoline: Model Evaluation under Distribution ShiftMayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms et al.ICML 2021 · 84 citations
- Estimating Generalization under Distribution Shifts via Domain-Invariant RepresentationsChing-Yao Chuang, Antonio Torralba, Stefanie JegelkaICML 2020 · 72 citations
Related papers
- Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2021 · 3 citations
- DAAP: Privacy-Preserving Model Accuracy Estimation on Unlabeled Datasets Through Distribution-Aware Adversarial PerturbationGuodong Cao, Zhibo Wang, Yunhe Feng, Xiaowei DongUSENIX Security 2024
- Density-driven Regularization for Out-of-distribution DetectionWenjian Huang, Hao Wang, Jiahao Xia, Chengyan Wang et al.NeurIPS 2022 · 17 citations
- (Almost) Provable Error Bounds Under Distribution Shift via Disagreement DiscrepancyElan Rosenfeld, Saurabh GargNeurIPS 2023 · 18 citations
- Unsupervised Accuracy Estimation of Deep Visual Models using Domain-Adaptive Adversarial Perturbation without Source SamplesJoonHo Lee, Jae Oh Woo, Hankyu Moon, Kwonho LeeICCV 2023 · 5 citations
