A Call to Reflect on Evaluation Practices for Failure Detection in Image Classification
Paul F. Jaeger, Carsten T. Lüth, Lukas Klein, Till J. Bungert
Abstract
Reliable application of machine learning-based decision systems in the wild is one of the major challenges currently investigated by the field. A large portion of established approaches aims to detect erroneous predictions by means of assigning confidence scores. This confidence may be obtained by either quantifying the model's predictive uncertainty, learning explicit scoring functions, or assessing whether the input is in line with the training distribution. Curiously, while these approaches all state to address the same eventual goal of detecting failures of a classifier upon real-world application, they currently constitute largely separated research fields with individual evaluation protocols, which either exclude a substantial part of relevant methods or ignore large parts of relevant failure sources. In this work, we systematically reveal current pitfalls caused by these inconsistencies and derive requirements for a holistic and realistic evaluation of failure detection. To demonstrate the relevance of this unified perspective, we present a large-scale empirical study for the first time enabling benchmarking confidence scoring functions w.r.t. all relevant methods and failure sources. The revelation of a simple softmax response baseline as the overall best performing method underlines the drastic shortcomings of current evaluation in the abundance of publicized research on confidence scoring. Code and trained models are at https://github.com/IML-DKFZ/fd-shifts .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63b00528-a819-4d40-9cca-e66472f1dc27Cited by top-tier papers21
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz et al.ICCV 2023 · 130 citations
- Overcoming Common Flaws in the Evaluation of Selective Classification SystemsJeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner et al.NeurIPS 2024 · 44 citations
- Navigating the Pitfalls of Active Learning Evaluation: A Systematic Framework for Meaningful Performance AssessmentCarsten T. Lüth, Till J. Bungert, Lukas Klein, Paul F. JaegerNeurIPS 2023 · 32 citations
- ValUES: A Framework for Systematic Validation of Uncertainty Estimation in Semantic SegmentationKim-Celine Kahl, Carsten T. Lüth, Maximilian Zenk, Klaus H. Maier-Hein et al.ICLR 2024 · 28 citations
- ImageNet-OOD: Deciphering Modern Out-of-Distribution Detection AlgorithmsWilliam Yang, Byron Zhang, Olga RussakovskyICLR 2024 · 23 citations
Builds on6
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- A Fine-Grained Analysis on Distribution ShiftOlivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi et al.ICLR 2022 · 258 citations
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 193 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 103 citations
Related papers
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Breaking Down Out-of-Distribution Detection: Many Methods Based on OOD Training Data Estimate a Combination of the Same Core QuantitiesJulian Bitterwolf, Alexander Meinke, Maximilian Augustin, Matthias HeinICML 2022 · 35 citations
- Beyond calibration: estimating the grouping loss of modern neural networksAlexandre Perez-Lebel, Marine Le Morvan, Gaël VaroquauxICLR 2023 · 5 citations
- Interpretable Failure Detection with Human-Level ConceptsKien X. Nguyen, Tang Li, Xi PengAAAI 2025 · 3 citations
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.NeurIPS 2024 · 13 citations
