Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification
Takashi Ishida, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu, Masashi Sugiyama
Abstract
There is a fundamental limitation in the prediction performance that a machine learning model can achieve due to the inevitable uncertainty of the prediction target. In classification problems, this can be characterized by the Bayes error, which is the best achievable error with any classifier. The Bayes error can be used as a criterion to evaluate classifiers with state-of-the-art performance and can be used to detect test set overfitting. We propose a simple and direct Bayes error estimator, where we just take the mean of the labels that show uncertainty of the class assignments. Our flexible approach enables us to perform Bayes error estimation even for weakly supervised data. In contrast to others, our method is model-free and even instancefree. Moreover, it has no hyperparameters and gives a more accurate estimate of the Bayes error than several baselines empirically. Experiments using our method suggest that recently proposed deep networks such as the Vision Transformer may have reached, or is about to reach, the Bayes error for benchmark datasets. Finally, we discuss how we can study the inherent difficulty of the acceptance/rejection decision for scientific articles, by estimating the Bayes error of the ICLR papers from 2017 to 2023. * Currently at Preferred Networks. 1 AUC stands for "Area under the ROC Curve." 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b9e8a4c-62cf-4398-82a2-88deaf768598Cited by top-tier papers8
- Binary Classification with Confidence DifferenceWei Wang, Lei Feng, Yuchen Jiang, Gang Niu et al.NeurIPS 2023 · 20 citations
- Demystifying the Optimal Performance of Multi-Class ClassificationMinoh Jeong, Martina Cardone, Alex DytsoNeurIPS 2023 · 17 citations
- Beyond probability partitions: Calibrating neural networks with semantic aware groupingJia-Qi Yang, De-Chuan Zhan, Le GanNeurIPS 2023 · 14 citations
- A General Framework for Learning from Weak SupervisionHao Chen, Jindong Wang, Lei Feng, Xiang Li et al.ICML 2024 · 13 citations
- Practical estimation of the optimal classification error with soft labels and calibrationRyota Ushio, Takashi Ishida, Masashi SugiyamaICLR 2026 · 6 citations
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- Evaluating Machine Accuracy on ImageNetVaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang et al.ICML 2020 · 153 citations
- Evaluating State-of-the-Art Classification Models Against Bayes OptimalityRyan Theisen, Huan Wang, Lav R. Varshney, Caiming Xiong et al.NeurIPS 2021 · 21 citations
Related papers
- Estimating Generalization under Distribution Shifts via Domain-Invariant RepresentationsChing-Yao Chuang, Antonio Torralba, Stefanie JegelkaICML 2020 · 72 citations
- Active Bayesian Assessment of Black-Box ClassifiersDisi Ji, Robert L. Logan IV, Padhraic Smyth, Mark SteyversAAAI 2021 · 3 citations
- Automatic Feasibility Study via Data Quality Analysis for ML: A Case-Study on Label NoiseCédric Renggli, Luka Rimanic, Luka Kolar, Wentao Wu et al.ICDE 2023 · 8 citations
- Conformal C2ST: Turning weak classifiers into strong two-sample testsVansh Bansal, Tianyu Chen, James ScottICML 2026 · 1 citation
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu et al.NeurIPS 2021 · 79 citations
