Interpretable Failure Detection with Human-Level Concepts
Kien X. Nguyen, Tang Li, Xi Peng
Abstract
Reliable failure detection holds paramount importance in safety-critical applications. Yet, neural networks are known to produce overconfident predictions for misclassified samples. As a result, it remains a problematic matter as existing confidence score functions rely on category-level signals, the logits, to detect failures. This research introduces an innovative strategy, leveraging human-level concepts for a dual purpose: to reliably detect when a model fails and to transparently interpret why. By integrating a nuanced array of signals for each category, our method enables a finer-grained assessment of the model's confidence. We present a simple yet highly effective approach based on the ordinal ranking of concept activation to the input image. Without bells and whistles, our method is able to significantly reduce the false positive rate across diverse real-world image classification benchmarks, specifically by 3.7% on ImageNet and 9.0% on EuroSAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba2d6d51-10ce-4bac-b37a-a85e990e405bCited by top-tier papers2
- Adaptive Confidence Regularization for Multimodal Failure DetectionMoru Liu, Hao Dong, Olga Fink, Mario TrappCVPR 2026 · 1 citation
- Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision TransformersTang Li, Yanlin Chen, Mengmeng Ma, Xi PengICML 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- ReAct: Out-of-distribution Detection With Rectified ActivationsYiyou Sun, Chuan Guo, Yixuan LiNeurIPS 2021 · 733 citations
- Confidence-Aware Learning for Deep Neural NetworksJooyoung Moon, Jihyo Kim, Younghak Shin, Sangheum HwangICML 2020 · 184 citations
- PRIME: Prioritizing Interpretability in Failure Mode ExtractionKeivan Rezaei, Mehrdad Saberi, Mazda Moayeri, Soheil FeiziICLR 2024 · 9 citations
- A Call to Reflect on Evaluation Practices for Failure Detection in Image ClassificationPaul F. Jaeger, Carsten T. Lüth, Lukas Klein, Till J. BungertICLR 2023 · 18 citations
- CE-FAM: Concept-Based Explanation via Fusion of Activation MapsMichihiro Kuroki, Toshihiko YamasakiICCV 2025 · 3 citations
