A Call to Reflect on Evaluation Practices for Failure Detection in Image Classification
Paul F. Jaeger, Carsten T. Lüth, Lukas Klein, Till J. Bungert
摘要
Reliable application of machine learning-based decision systems in the wild is one of the major challenges currently investigated by the field. A large portion of established approaches aims to detect erroneous predictions by means of assigning confidence scores. This confidence may be obtained by either quantifying the model's predictive uncertainty, learning explicit scoring functions, or assessing whether the input is in line with the training distribution. Curiously, while these approaches all state to address the same eventual goal of detecting failures of a classifier upon real-world application, they currently constitute largely separated research fields with individual evaluation protocols, which either exclude a substantial part of relevant methods or ignore large parts of relevant failure sources. In this work, we systematically reveal current pitfalls caused by these inconsistencies and derive requirements for a holistic and realistic evaluation of failure detection. To demonstrate the relevance of this unified perspective, we present a large-scale empirical study for the first time enabling benchmarking confidence scoring functions w.r.t. all relevant methods and failure sources. The revelation of a simple softmax response baseline as the overall best performing method underlines the drastic shortcomings of current evaluation in the abundance of publicized research on confidence scoring. Code and trained models are at https://github.com/IML-DKFZ/fd-shifts .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz 等ICCV 2023 · 被引用 130 次
- Overcoming Common Flaws in the Evaluation of Selective Classification SystemsJeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner 等NeurIPS 2024 · 被引用 44 次
- Navigating the Pitfalls of Active Learning Evaluation: A Systematic Framework for Meaningful Performance AssessmentCarsten T. Lüth, Till J. Bungert, Lukas Klein, Paul F. JaegerNeurIPS 2023 · 被引用 32 次
- ValUES: A Framework for Systematic Validation of Uncertainty Estimation in Semantic SegmentationKim-Celine Kahl, Carsten T. Lüth, Maximilian Zenk, Klaus H. Maier-Hein 等ICLR 2024 · 被引用 28 次
- ImageNet-OOD: Deciphering Modern Out-of-Distribution Detection AlgorithmsWilliam Yang, Byron Zhang, Olga RussakovskyICLR 2024 · 被引用 23 次
它引用的顶会 Paper6
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- A Fine-Grained Analysis on Distribution ShiftOlivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi 等ICLR 2022 · 被引用 258 次
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 被引用 193 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 被引用 103 次
相关 Paper
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell 等ICCV 2021 · 被引用 141 次
- Breaking Down Out-of-Distribution Detection: Many Methods Based on OOD Training Data Estimate a Combination of the Same Core QuantitiesJulian Bitterwolf, Alexander Meinke, Maximilian Augustin, Matthias HeinICML 2022 · 被引用 35 次
- Beyond calibration: estimating the grouping loss of modern neural networksAlexandre Perez-Lebel, Marine Le Morvan, Gaël VaroquauxICLR 2023 · 被引用 5 次
- Interpretable Failure Detection with Human-Level ConceptsKien X. Nguyen, Tang Li, Xi PengAAAI 2025 · 被引用 3 次
- Sequential Harmful Shift Detection Without LabelsSalim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
