Grounding Representation Similarity Through Statistical Testing
Frances Ding, Jean-Stanislas Denain, Jacob Steinhardt
Abstract
To understand neural network behavior, recent works quantitatively compare different networks' learned representations using canonical correlation analysis (CCA), centered kernel alignment (CKA), and other dissimilarity measures. Unfortunately, these widely used measures often disagree on fundamental observations, such as whether deep networks differing only in random initialization learn similar representations. These disagreements raise the question: which, if any, of these dissimilarity measures should we believe? We provide a framework to ground this question through a concrete test: measures should have sensitivity to changes that affect functional behavior, and specificity against changes that do not. We quantify this through a variety of functional behaviors including probing accuracy and robustness to distribution shift, and examine changes such as varying random initialization and deleting principal components. We find that current metrics exhibit different weaknesses, note that a classical baseline performs surprisingly well, and highlight settings where all metrics appear to fail, thus providing a challenge set for further improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d35a4116-ba54-49ef-8c6d-4b82d231aea1Cited by top-tier papers22
- On the Symmetries of Deep Learning Models and their Internal RepresentationsCharles Godfrey, Davis Brown, Tegan Emerson, Henry KvingeNeurIPS 2022 · 78 citations
- Architecture Agnostic Federated Learning for Neural NetworksDisha Makhija, Xing Han, Nhat Ho, Joydeep GhoshICML 2022 · 62 citations
- Dataset Inference for Self-Supervised ModelsAdam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan et al.NeurIPS 2022 · 59 citations
- Revisiting the Platonic Representation Hypothesis: An Aristotelian ViewFabian Gröger, Shuo Wen, Maria BrbicICML 2026 · 27 citations
- Experimental Observations of the Topology of Convolutional Neural Network ActivationsEmilie Purvine, Davis Brown, Brett A. Jefferson, Cliff A. Joslyn et al.AAAI 2023 · 21 citations
Builds on4
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Knowledge Consistency between Neural Networks and BeyondRuofan Liang, Tianlin Li, Longfei Li, Jing Wang et al.ICLR 2020 · 30 citations
Related papers
- Reliability of CKA as a Similarity Measure in Deep LearningMohammadReza Davari, Stefan Horoi, Amine Natik, Guillaume Lajoie et al.ICLR 2023 · 3 citations
- Generalized Shape Metrics on Neural RepresentationsAlex H. Williams, Erin Kunz, Simon Kornblith, Scott W. LindermanNeurIPS 2021 · 182 citations
- Differentiable Optimization of Similarity Scores Between Models and BrainsNathan Cloos, Moufan Li, Markus Siegel, Scott L. Brincat et al.ICLR 2025 · 1 citation
- Deconfounded Representation Similarity for Comparison of Neural NetworksTianyu Cui, Yogesh Kumar, Pekka Marttinen, Samuel KaskiNeurIPS 2022 · 27 citations
- Spectral Analysis of Representational Similarity with Limited NeuronsHyunmo Kang, Abdulkadir Canatar, SueYeon ChungNeurIPS 2025 · 4 citations
