Optimal Binary Classification Beyond Accuracy
Shashank Singh, Justin T. Khim
Abstract
The vast majority of statistical theory on binary classification characterizes performance in terms of accuracy. However, accuracy is known in many cases to poorly reflect the practical consequences of classification error, most famously in imbalanced binary classification, where data are dominated by samples from one of two classes. The first part of this paper derives a novel generalization of the Bayes-optimal classifier from accuracy to any performance metric computed from the confusion matrix. Specifically, this result (a) demonstrates that stochastic classifiers sometimes outperform the best possible deterministic classifier and (b) removes an empirically unverifiable absolute continuity assumption that is poorly understood but pervades existing results. We then demonstrate how to use this generalized Bayes classifier to obtain regret bounds in terms of the error of estimating regression functions under uniform loss. Finally, we use these results to develop some of the first finite-sample statistical guarantees specific to imbalanced binary classification. Specifically, we demonstrate that optimal classification performance depends on properties of class imbalance, such as a novel notion called Uniform Class Imbalance, that have not previously been formalized. We further illustrate these contributions numerically in the case of k-nearest neighbor classification. ˚The contributions in this paper were made prior to joining Amazon. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bd8bff5-877b-4d13-a6f6-b352ccd4d245Cited by top-tier papers5
- On Missing Labels, Long-tails and Propensities in Extreme Multi-label ClassificationErik Schultheis, Marek Wydmuch, Rohit Babbar, Krzysztof DembczynskiKDD 2022 · 20 citations
- How Far Can Fairness Constraints Help Recover From Biased Data?Mohit Sharma, Amit DeshpandeICML 2024 · 7 citations
- Consistent algorithms for multi-label classification with macro-at-k metricsErik Schultheis, Wojciech Kotlowski, Marek Wydmuch, Rohit Babbar et al.ICLR 2024 · 6 citations
- On Optimal Steering to Achieve Exact FairnessMohit Sharma, Amit Deshpande, Chiranjib Bhattacharyya, Rajiv Ratn ShahNeurIPS 2025 · 2 citations
- A General Online Algorithm for Optimizing Complex Performance MetricsWojciech Kotlowski, Marek Wydmuch, Erik Schultheis, Rohit Babbar et al.ICML 2024 · 1 citation
Builds on3
Related papers
- Learning to Reject Meets Long-tail LearningHarikrishna Narasimhan, Aditya Krishna Menon, Wittawat Jitkrittum, Neha Gupta et al.ICLR 2024 · 7 citations
- PAC-Bayes Analysis for Recalibration in ClassificationMasahiro Fujisawa, Futoshi FutamiICML 2025
- Improved Algorithms for Neural Active LearningYikun Ban, Yuheng Zhang, Hanghang Tong, Arindam Banerjee et al.NeurIPS 2022 · 18 citations
- Regret Bounds for Multilabel Classification in Sparse Label RegimesRóbert Busa-Fekete, Heejin Choi, Krzysztof Dembczynski, Claudio Gentile et al.NeurIPS 2022 · 5 citations
- Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced DataCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
