Overcoming Common Flaws in the Evaluation of Selective Classification Systems
Jeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner, Klaus H. Maier-Hein, Lena Maier-Hein, Paul F. Jaeger
摘要
Selective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of these systems typically assumes fixed working points based on pre-defined rejection thresholds, methodological progress requires benchmarking the general performance of systems akin to the in standard classification. In this work, we define 5 requirements for multi-threshold metrics in selective classification regarding task alignment, interpretability, and flexibility, and show how current approaches fail to meet them. We propose the Area under the Generalized Risk Coverage curve (), which meets all requirements and can be directly interpreted as the average risk of undetected failures. We empirically demonstrate the relevance of on a comprehensive benchmark spanning 6 data sets and 13 confidence scoring functions. We find that the proposed metric substantially changes metric rankings on 5 out of the 6 data sets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What Does It Take to Build a Performant Selective Classifier?Stephan Rabanser, Nicolas PapernotNeurIPS 2025 · 被引用 7 次
- Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream PerformanceVanessa Emanuela Guarino, Claudia Winklmayr, Jannik Franzen, Josef Rumberger 等CVPR 2026 · 被引用 1 次
- Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective PredictionAditya Sarkar, Yi Li, Jiacheng Cheng, Shlok Kumar Mishra 等ICLR 2026
- A Novel Characterization of the Population Area Under the Risk Coverage Curve (AURC) and Rates of Finite Sample EstimatorsHan Zhou, Jordy Van Landeghem, Teodora Popordanoska, Matthew B. BlaschkoICML 2025
- Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix ItGuoxuan Xia, Olivier Laurent, Gianni Franchi, Christos-Savvas BouganisICLR 2025
它引用的顶会 Paper10
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 被引用 354 次
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz 等ICCV 2023 · 被引用 130 次
- Disrupting Deep Uncertainty Estimation Without Harming AccuracyIdo Galil, Ran El-YanivNeurIPS 2021 · 被引用 27 次
相关 Paper
- Towards Decision-Friendly AUC: Learning Multi-Classifier with AUCµPeifeng Gao, Qianqian Xu, Peisong Wen, Huiyang Shao 等AAAI 2023 · 被引用 1 次
- The VOROS: Lifting ROC Curves to 3D to Summarize Unbalanced Classifier PerformanceChristopher Ratigan, Lenore CowenAAAI 2025 · 被引用 1 次
- AUC Optimization with a Reject OptionSong-Qing Shen, Bin-Bin Yang, Wei GaoAAAI 2020 · 被引用 7 次
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen 等AAAI 2026
- Conditional Coverage Diagnostics for Conformal PredictionSacha Braun, David Holzmüller, Michael Jordan, Francis BachICML 2026 · 被引用 12 次
