Overinterpretation reveals image classification model pathologies
Brandon Carter, Siddhartha Jain, Jonas Mueller, David Gifford
Abstract
Image classifiers are typically scored on their test set accuracy, but high accuracy can mask a subtle type of model failure. We find that high scoring convolutional neural networks (CNNs) on popular benchmarks exhibit troubling pathologies that allow them to display high accuracy even in the absence of semantically salient features. When a model provides a high-confidence decision without salient supporting input features, we say the classifier has overinterpreted its input, finding too much class-evidence in patterns that appear nonsensical to humans. Here, we demonstrate that neural networks trained on CIFAR-10 and ImageNet suffer from overinterpretation, and we find models on CIFAR-10 make confident predictions even when 95% of input images are masked and humans cannot discern salient features in the remaining pixel-subsets. We introduce Batched Gradient SIS, a new method for discovering sufficient input subsets for complex datasets, and use this method to show the sufficiency of border pixels in ImageNet for training and testing. Although these patterns portend potential model fragility in real-world deployment, they are in fact valid statistical patterns of the benchmark that alone suffice to attain high test accuracy. Unlike adversarial examples, overinterpretation relies upon unmodified image pixels. We find ensembling and input dropout can each help mitigate overinterpretation. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c58053a9-dbb0-4bf2-8dd6-4a2051d06d84Cited by top-tier papers11
- Optimizing Relevance Maps of Vision Transformers Improves RobustnessHila Chefer, Idan Schwartz, Lior WolfNeurIPS 2022 · 55 citations
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt et al.ICCV 2025 · 47 citations
- Missingness Bias in Model DebuggingSaachi Jain, Hadi Salman, Eric Wong, Pengchuan Zhang et al.ICLR 2022 · 45 citations
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 43 citations
Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Saliency is a Possible Red Herring When Diagnosing Poor GeneralizationJoseph D. Viviano, Becks Simpson, Francis Dutil, Yoshua Bengio et al.ICLR 2021 · 46 citations
Related papers
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- DANCE: Enhancing saliency maps using decoysYang Young Lu, Wenbo Guo, Xinyu Xing, William Stafford NobleICML 2021 · 14 citations
- Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNetYannic Neuhaus, Maximilian Augustin, Valentyn Boreiko, Matthias HeinICCV 2023 · 42 citations
- From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature SelectionMoritz Vandenhirtz, Julia E. VogtICML 2025
- AIM: Amending Inherent Interpretability via Self-Supervised MaskingEyad Alshami, Shashank Agnihotri, Bernt Schiele, Margret KeuperICCV 2025 · 3 citations
