Optimal Binary Classification Beyond Accuracy
Shashank Singh, Justin T. Khim
摘要
The vast majority of statistical theory on binary classification characterizes performance in terms of accuracy. However, accuracy is known in many cases to poorly reflect the practical consequences of classification error, most famously in imbalanced binary classification, where data are dominated by samples from one of two classes. The first part of this paper derives a novel generalization of the Bayes-optimal classifier from accuracy to any performance metric computed from the confusion matrix. Specifically, this result (a) demonstrates that stochastic classifiers sometimes outperform the best possible deterministic classifier and (b) removes an empirically unverifiable absolute continuity assumption that is poorly understood but pervades existing results. We then demonstrate how to use this generalized Bayes classifier to obtain regret bounds in terms of the error of estimating regression functions under uniform loss. Finally, we use these results to develop some of the first finite-sample statistical guarantees specific to imbalanced binary classification. Specifically, we demonstrate that optimal classification performance depends on properties of class imbalance, such as a novel notion called Uniform Class Imbalance, that have not previously been formalized. We further illustrate these contributions numerically in the case of k-nearest neighbor classification. ˚The contributions in this paper were made prior to joining Amazon. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On Missing Labels, Long-tails and Propensities in Extreme Multi-label ClassificationErik Schultheis, Marek Wydmuch, Rohit Babbar, Krzysztof DembczynskiKDD 2022 · 被引用 20 次
- How Far Can Fairness Constraints Help Recover From Biased Data?Mohit Sharma, Amit DeshpandeICML 2024 · 被引用 7 次
- Consistent algorithms for multi-label classification with macro-at-k metricsErik Schultheis, Wojciech Kotlowski, Marek Wydmuch, Rohit Babbar 等ICLR 2024 · 被引用 6 次
- On Optimal Steering to Achieve Exact FairnessMohit Sharma, Amit Deshpande, Chiranjib Bhattacharyya, Rajiv Ratn ShahNeurIPS 2025 · 被引用 2 次
- A General Online Algorithm for Optimizing Complex Performance MetricsWojciech Kotlowski, Marek Wydmuch, Erik Schultheis, Rohit Babbar 等ICML 2024 · 被引用 1 次
它引用的顶会 Paper3
相关 Paper
- Learning to Reject Meets Long-tail LearningHarikrishna Narasimhan, Aditya Krishna Menon, Wittawat Jitkrittum, Neha Gupta 等ICLR 2024 · 被引用 7 次
- PAC-Bayes Analysis for Recalibration in ClassificationMasahiro Fujisawa, Futoshi FutamiICML 2025
- Improved Algorithms for Neural Active LearningYikun Ban, Yuheng Zhang, Hanghang Tong, Arindam Banerjee 等NeurIPS 2022 · 被引用 18 次
- Regret Bounds for Multilabel Classification in Sparse Label RegimesRóbert Busa-Fekete, Heejin Choi, Krzysztof Dembczynski, Claudio Gentile 等NeurIPS 2022 · 被引用 5 次
- Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced DataCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
