Decoupling Semantic Similarity from Spatial Alignment for Neural Networks
Tassilo Wald, Constantin Ulrich, Priyank Jaini, Gregor Köhler, David Zimmerer, Stefan Denner, Fabian Isensee, Michael Baumgartner, Klaus H. Maier-Hein
Abstract
What representation do deep neural networks learn? How similar are images to each other for neural networks? Despite the overwhelming success of deep learning methods key questions about their internal workings still remain largely unanswered, due to their internal high dimensionality and complexity. To address this, one approach is to measure the similarity of activation responses to various inputs. Representational Similarity Matrices (RSMs) distill this similarity into scalar values for each input pair. These matrices encapsulate the entire similarity structure of a system, indicating which input leads to similar responses. While the similarity between images is ambiguous, we argue that the spatial location of semantic objects does neither influence human perception nor deep learning classifiers. Thus this should be reflected in the definition of similarity between image responses for computer vision systems. Revisiting the established similarity calculations for RSMs we expose their sensitivity to spatial alignment. In this paper, we propose to solve this through semantic RSMs, which are invariant to spatial permutation. We measure semantic similarity between input responses by formulating it as a set-matching problem. Further, we quantify the superiority of semantic RSMs over spatio-semantic RSMs through image retrieval and by comparing the similarity between representations to the similarity between predicted class probabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04e4c52a-7f2b-49b8-8ebe-e5ba4aa31638Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
Related papers
- Generalized Shape Metrics on Neural RepresentationsAlex H. Williams, Erin Kunz, Simon Kornblith, Scott W. LindermanNeurIPS 2021 · 182 citations
- Representational Similarity via Interpretable Visual ConceptsNeehar Kondapaneni, Oisin Mac Aodha, Pietro PeronaICLR 2025
- Deconfounded Representation Similarity for Comparison of Neural NetworksTianyu Cui, Yogesh Kumar, Pekka Marttinen, Samuel KaskiNeurIPS 2022 · 27 citations
- Towards Interpretable Deep Metric Learning with Structural MatchingWenliang Zhao, Yongming Rao, Ziyi Wang, Jiwen Lu et al.ICCV 2021 · 52 citations
- How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy ModelUmberto M. Tomasini, Matthieu WyartICML 2024 · 7 citations
