Overcoming Common Flaws in the Evaluation of Selective Classification Systems
Jeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner, Klaus H. Maier-Hein, Lena Maier-Hein, Paul F. Jaeger
Abstract
Selective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of these systems typically assumes fixed working points based on pre-defined rejection thresholds, methodological progress requires benchmarking the general performance of systems akin to the in standard classification. In this work, we define 5 requirements for multi-threshold metrics in selective classification regarding task alignment, interpretability, and flexibility, and show how current approaches fail to meet them. We propose the Area under the Generalized Risk Coverage curve (), which meets all requirements and can be directly interpreted as the average risk of undetected failures. We empirically demonstrate the relevance of on a comprehensive benchmark spanning 6 data sets and 13 confidence scoring functions. We find that the proposed metric substantially changes metric rankings on 5 out of the 6 data sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ece9ed5-ef66-416a-aefc-296d1b48c63cCited by top-tier papers5
- What Does It Take to Build a Performant Selective Classifier?Stephan Rabanser, Nicolas PapernotNeurIPS 2025 · 7 citations
- Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream PerformanceVanessa Emanuela Guarino, Claudia Winklmayr, Jannik Franzen, Josef Rumberger et al.CVPR 2026 · 1 citation
- Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective PredictionAditya Sarkar, Yi Li, Jiacheng Cheng, Shlok Kumar Mishra et al.ICLR 2026
- A Novel Characterization of the Population Area Under the Risk Coverage Curve (AURC) and Rates of Finite Sample EstimatorsHan Zhou, Jordy Van Landeghem, Teodora Popordanoska, Matthew B. BlaschkoICML 2025
- Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix ItGuoxuan Xia, Olivier Laurent, Gianni Franchi, Christos-Savvas BouganisICLR 2025
Builds on10
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz et al.ICCV 2023 · 130 citations
- Disrupting Deep Uncertainty Estimation Without Harming AccuracyIdo Galil, Ran El-YanivNeurIPS 2021 · 27 citations
Related papers
- Towards Decision-Friendly AUC: Learning Multi-Classifier with AUCµPeifeng Gao, Qianqian Xu, Peisong Wen, Huiyang Shao et al.AAAI 2023 · 1 citation
- The VOROS: Lifting ROC Curves to 3D to Summarize Unbalanced Classifier PerformanceChristopher Ratigan, Lenore CowenAAAI 2025 · 1 citation
- AUC Optimization with a Reject OptionSong-Qing Shen, Bin-Bin Yang, Wei GaoAAAI 2020 · 7 citations
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen et al.AAAI 2026
- Conditional Coverage Diagnostics for Conformal PredictionSacha Braun, David Holzmüller, Michael Jordan, Francis BachICML 2026 · 12 citations
