When Random Saliency Looks Trained: Architectural Center Bias in CNN Interpretability
Keying Kuang, Iain Carmichael, Elizabeth Purdom
Abstract
Saliency maps are widely used to interpret image classification models and build trust in their predictions; however, their reliability remains a central concern, as randomized networks can produce saliency maps that closely resemble those of trained models. We identify a previously underappreciated architectural contributor to this phenomenon: a center-focused saliency bias induced by common convolutional design choices. Through controlled ablations, we show that architectural components such as zero padding and receptive field growth induce a center-focused saliency prior that is already present in randomly initialized CNNs and under randomized inputs. In contrast, this behavior is largely absent in non-convolutional architectures such as Vision Transformers (ViTs) and multilayer perceptrons (MLPs). To investigate the interaction between architectural priors and learning, we introduce a corner-shift benchmark and a Center-Shift Index that quantify how saliency redistributes under object relocation. We show that training can partially shift saliency toward object regions, while randomized models remain dominated by architectural center bias, providing one mechanism by which trained-random similarity can be inflated and clarifying how architectural priors can confound standard saliency evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 147 citations
- Mind the Pad - CNNs Can Develop Blind SpotsBilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan et al.ICLR 2021 · 32 citations
- Shortcomings of Top-Down Randomization-Based Sanity Checks for Evaluations of Deep Neural Network ExplanationsAlexander Binder, Leander Weber, Sebastian Lapuschkin, Grégoire Montavon et al.CVPR 2023
- On Translation Invariance in CNNs: Convolutional Layers Can Exploit Absolute Spatial LocationOsman Semih Kayhan, Jan C. van GemertCVPR 2020
Related papers
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 115 citations
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
- Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models Against Gaussian NoiseBum Jun Kim, Makoto Kawano, Yusuke Iwasawa, Yutaka MatsuoICML 2026
- Alias-Free ViT: Fractional Shift Invariance via Linear AttentionHagay Michaeli, Daniel SoudryNeurIPS 2025 · 2 citations
- DAVE: Distribution-aware Attribution via ViT Gradient DecompositionAdam Wróbel, Siddhartha Gairola, Jacek Tabor, Bernt Schiele et al.ICML 2026 · 2 citations
