I Am Big, You Are Little; I Am Right, You Are Wrong
David A. Kelly, Akchunya Chanchal, Nathan Blake
Abstract
Machine learning for image classification is an active and rapidly developing field. With the proliferation of classifiers of different sizes and different architectures, the problem of choosing the right model becomes more and more important. While we can assess a model's classification accuracy statistically, our understanding of the way these models work is unfortunately limited. In order to gain insight into the decision-making process of different vision models, we propose using minimal sufficient pixels sets to gauge a model's ‘concentration’: the pixels that capture the essence of an image through the lens of the model. By comparing position, overlap, and size of sets of pixels, we identify that different architectures have statistically different concentration, in both size and position. In particular, ConvNext and EVA models differ markedly from the others. We also identify that images which are misclassified are associated with larger pixels sets than correct classifications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4bb2207-eb59-4b5d-b668-973784a5dcd4Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- When does dough become a bagel? Analyzing the remaining mistakes on ImageNetVijay Vasudevan, Benjamin Caine, Raphael Gontijo Lopes, Sara Fridovich-Keil et al.NeurIPS 2022 · 79 citations
Related papers
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Understanding Intrinsic Robustness Using Label UncertaintyXiao Zhang, David E. EvansICLR 2022 · 6 citations
- Consensus vs. Controversy: Mapping the Decision Space Where Architectures DivergeMinhyeok LeeCVPR 2026
- Towards Disentangling Information Paths with Coded ResNeXtApostolos Avranas, Marios KountourisNeurIPS 2022 · 1 citation
- Identifying Important Group of Pixels using InteractionsKosuke Sumiyasu, Kazuhiko Kawamoto, Hiroshi KeraCVPR 2024
