Comparing the Decision-Making Mechanisms by Transformers and CNNs via Explanation Methods
Mingqi Jiang, Saeed Khorram, Fuxin Li
Abstract
In order to gain insights about the decision-making of different visual recognition backbones, we propose two methodologies, sub-explanation counting and cross-testing, that systematically applies deep explanation algorithms on a dataset-wide basis, and compares the statistics generated from the amount and nature of the explanations. These methodologies reveal the difference among networks in terms of two properties called compositionality and disjunctivism. Transformers and ConvNeXt are found to be more compositional, in the sense that they jointly consider multiple parts of the image in building their decisions, whereas traditional CNNs and distilled transformers are less compositional and more disjunctive, which means that they use multiple diverse but smaller set of parts to achieve a confident prediction. Through further experiments, we pinpointed the choice of normalization to be especially important in the compositionality of a model, in that batch normalization leads to less compositionality while group and layer normalization lead to more. Finally, we also analyze the features shared by different backbones and plot a landscape of different models based on their feature-use similarity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- I Am Big, You Are Little; I Am Right, You Are WrongDavid A. Kelly, Akchunya Chanchal, Nathan BlakeICCV 2025 · 9 citations
- PhaseWin Search Framework Enable Efficient Object-Level InterpretationZihan Gu, Ruoyu Chen, Junchi Zhang, Yue Hu et al.CVPR 2026 · 1 citation
- Interpreting Object-level Foundation Models via Visual Precision SearchRuoyu Chen, Siyuan Liang, Jingzhi Li, Shiming Liu et al.CVPR 2025
- Explaining Object Detectors via Collective Contribution of PixelsToshinori Yamauchi, Hiroshi Kera, Kazuhiko KawamotoCVPR 2026
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Labeling Neural Representations with Inverse RecognitionKirill Bykov, Laura Kopf, Shinichi Nakajima, Marius Kloft et al.NeurIPS 2023 · 36 citations
- Consensus vs. Controversy: Mapping the Decision Space Where Architectures DivergeMinhyeok LeeCVPR 2026
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsMichael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre et al.NeurIPS 2024 · 22 citations
- DISSECT: Disentangled Simultaneous Explanations via Concept TraversalsAsma Ghandeharioun, Been Kim, Chun-Liang Li, Brendan Jou et al.ICLR 2022 · 58 citations
