Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification
Takashi Ishida, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu, Masashi Sugiyama
摘要
There is a fundamental limitation in the prediction performance that a machine learning model can achieve due to the inevitable uncertainty of the prediction target. In classification problems, this can be characterized by the Bayes error, which is the best achievable error with any classifier. The Bayes error can be used as a criterion to evaluate classifiers with state-of-the-art performance and can be used to detect test set overfitting. We propose a simple and direct Bayes error estimator, where we just take the mean of the labels that show uncertainty of the class assignments. Our flexible approach enables us to perform Bayes error estimation even for weakly supervised data. In contrast to others, our method is model-free and even instancefree. Moreover, it has no hyperparameters and gives a more accurate estimate of the Bayes error than several baselines empirically. Experiments using our method suggest that recently proposed deep networks such as the Vision Transformer may have reached, or is about to reach, the Bayes error for benchmark datasets. Finally, we discuss how we can study the inherent difficulty of the acceptance/rejection decision for scientific articles, by estimating the Bayes error of the ICLR papers from 2017 to 2023. * Currently at Preferred Networks. 1 AUC stands for "Area under the ROC Curve." 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Binary Classification with Confidence DifferenceWei Wang, Lei Feng, Yuchen Jiang, Gang Niu 等NeurIPS 2023 · 被引用 20 次
- Demystifying the Optimal Performance of Multi-Class ClassificationMinoh Jeong, Martina Cardone, Alex DytsoNeurIPS 2023 · 被引用 17 次
- Beyond probability partitions: Calibrating neural networks with semantic aware groupingJia-Qi Yang, De-Chuan Zhan, Le GanNeurIPS 2023 · 被引用 14 次
- A General Framework for Learning from Weak SupervisionHao Chen, Jindong Wang, Lei Feng, Xiang Li 等ICML 2024 · 被引用 13 次
- Practical estimation of the optimal classification error with soft labels and calibrationRyota Ushio, Takashi Ishida, Masashi SugiyamaICLR 2026 · 被引用 6 次
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- Evaluating Machine Accuracy on ImageNetVaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang 等ICML 2020 · 被引用 153 次
- Evaluating State-of-the-Art Classification Models Against Bayes OptimalityRyan Theisen, Huan Wang, Lav R. Varshney, Caiming Xiong 等NeurIPS 2021 · 被引用 21 次
相关 Paper
- Estimating Generalization under Distribution Shifts via Domain-Invariant RepresentationsChing-Yao Chuang, Antonio Torralba, Stefanie JegelkaICML 2020 · 被引用 72 次
- Active Bayesian Assessment of Black-Box ClassifiersDisi Ji, Robert L. Logan IV, Padhraic Smyth, Mark SteyversAAAI 2021 · 被引用 3 次
- Automatic Feasibility Study via Data Quality Analysis for ML: A Case-Study on Label NoiseCédric Renggli, Luka Rimanic, Luka Kolar, Wentao Wu 等ICDE 2023 · 被引用 8 次
- Conformal C2ST: Turning weak classifiers into strong two-sample testsVansh Bansal, Tianyu Chen, James ScottICML 2026 · 被引用 1 次
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu 等NeurIPS 2021 · 被引用 79 次
