AIM: Amending Inherent Interpretability via Self-Supervised Masking
Eyad Alshami, Shashank Agnihotri, Bernt Schiele, Margret Keuper
Abstract
It has been observed that deep neural networks (DNNs) often use both genuine as well as spurious features. In this work, we propose “Amending Inherent Interpretability via Self-Supervised Masking” (AIM), a simple yet interestingly effective method that promotes the network's utilization of genuine features over spurious alternatives without requiring additional annotations. In particular, AIM uses features at multiple encoding stages to guide a selfsupervised, sample-specific feature-masking process. As a result, AIM enables the training of well-performing and inherently interpretable models that faithfully summarize the decision process. We validate AIM across a diverse range of challenging datasets that test both out-of-distribution generalization and fine-grained visual understanding. These include general-purpose classification benchmarks such as ImageNet100, HardImageNet, and ImageWoof, as well as fine-grained classification datasets such as Waterbirds, TravelingBirds, and CUB-200. AIM demonstrates significant dual benefits: interpretability improvements, as measured by the Energy Pointing Game (EPG) score, and accuracy gains over strong baselines. These consistent gains across domains and architectures provide compelling evidence that AIM promotes the use of genuine and meaningful features that directly contribute to improved generalization and human-aligned interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5e30252-004c-4bc5-a89d-0d2a2fa3c44dBuilds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Harmonizing the object recognition strategies of deep neural networks with humansThomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas SerreNeurIPS 2022 · 111 citations
- MaskTune: Mitigating Spurious Correlations by Forcing to ExploreSaeid Asgari Taghanaki, Aliasghar Khani, Fereshte Khani, Ali Gholami et al.NeurIPS 2022 · 74 citations
Related papers
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- Enhancing Interpretability for Vision Models via Shapley Value OptimizationKanglong Fan, Yunqiao Yang, Chen MaAAAI 2026
- On Feature Learning in the Presence of Spurious CorrelationsPavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon WilsonNeurIPS 2022 · 208 citations
- Removing Spurious Concepts from Neural Network Representations via Joint Subspace EstimationFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICML 2024 · 3 citations
- Overcoming Simplicity Bias in Deep Networks using a Feature SieveRishabh Tiwari, Pradeep ShenoyICML 2023 · 32 citations
