Consensus vs. Controversy: Mapping the Decision Space Where Architectures Diverge
Minhyeok Lee
Abstract
Modern computer vision models from different architecture families, e.g., CNNs, Vision Transformers, and MLP-Mixers, achieve remarkably similar aggregate performance on standard benchmarks, masking potential systematic differences in how they process visual information. We introduce a simple yet revealing framework to identify where pretrained model families diverge: by systematically mapping the high-disagreement tail versus the lowdisagreement tail of the image distribution. Analyzing 12 pretrained models spanning three architecture families on ImageNet validation set, we discover that controversial images exhibit approximately 4.5× higher disagreement than consensus images (Controversy Score: 4.46). Despite mean accuracy around 80%, models show structured disagreement patterns: within-family agreement exceeds cross-family agreement, with CNNs and ViTs forming distinct clusters while MLPs show lower overall alignment. Crucially, only the top 10% most controversial images drive the majority of architectural divergence, constituting a small but informationally dense subset that reveals fundamental differences masked by aggregate metrics. Our analysis demonstrates that architectural choice matters most on this concentrated controversy space, providing researchers with actionable guidance for model selection and ensemble construction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
Related papers
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet AccuracyKirill Vishniakov, Zhiqiang Shen, Zhuang LiuICML 2024 · 26 citations
- Comparing the Decision-Making Mechanisms by Transformers and CNNs via Explanation MethodsMingqi Jiang, Saeed Khorram, Fuxin LiCVPR 2024
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li et al.ICCV 2021 · 501 citations
- Sequencer: Deep LSTM for Image ClassificationYuki Tatsunami, Masato TakiNeurIPS 2022 · 124 citations
