Practical estimation of the optimal classification error with soft labels and calibration
Ryota Ushio, Takashi Ishida, Masashi Sugiyama
Abstract
While the performance of machine learning systems has experienced significant improvement in recent years, relatively little attention has been paid to the fundamental question: to what extent can we improve our models? This paper provides a means of answering this question in the setting of binary classification, which is practical and theoretically supported. We extend a previous work that utilizes soft labels for estimating the Bayes error, the optimal error rate, in two important ways. First, we theoretically investigate the properties of the bias of the hard-label-based estimator discussed in the original work. We reveal that the decay rate of the bias is adaptive to how well the two class-conditional distributions are separated, and it can decay significantly faster than the previous result suggested as the number of hard labels per instance grows. Second, we tackle a more challenging problem setting: estimation with corrupted soft labels. One might be tempted to use calibrated soft labels instead of clean ones. However, we reveal that calibration guarantee is not enough, that is, even perfectly calibrated soft labels can result in a substantially inaccurate estimate. Then, we show that isotonic calibration can provide a statistically consistent estimator under an assumption weaker than that of the previous work. Our method is instance-free, i.e., we do not assume access to any input instances. This feature allows it to be adopted in practical scenarios where the instances are not available due to privacy issues. Experiments with synthetic and real-world datasets show the validity of our methods and theory. The code is available at https://github.com/RyotaUshio/bayes-error-estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e817822d-261b-43fd-9ec8-3630dbbdd2fbCited by top-tier papers1
Ask how each one uses itBuilds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- What Can We Learn from Collective Human Opinions on Natural Language Inference Data?Yixin Nie, Xiang Zhou, Mohit BansalEMNLP 2020 · 77 citations
- Calibrating Reasoning in Language Models with Internal ConsistencyZhihui Xie, Jizhou Guo, Tong Yu, Shuai LiNeurIPS 2024 · 37 citations
- Evaluating State-of-the-Art Classification Models Against Bayes OptimalityRyan Theisen, Huan Wang, Lav R. Varshney, Caiming Xiong et al.NeurIPS 2021 · 21 citations
Related papers
- Demystifying the Optimal Performance of Multi-Class ClassificationMinoh Jeong, Martina Cardone, Alex DytsoNeurIPS 2023 · 17 citations
- Resurfacing the Instance-only Dependent Label Noise Model through Loss CorrectionMustafa Enes Aydın, Maarten De Vos, Alexander BertrandICLR 2026
- Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary ClassificationTakashi Ishida, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu et al.ICLR 2023 · 3 citations
- Measuring Uncertainty CalibrationKamil Ciosek, Nicolò Felicioni, Sina Ghiassian, Juan Elenter Litwin et al.ICLR 2026 · 1 citation
- Information-theoretic Generalization Analysis for Expected Calibration ErrorFutoshi Futami, Masahiro FujisawaNeurIPS 2024 · 22 citations
