Revisiting Model Stitching to Compare Neural Representations
Yamini Bansal, Preetum Nakkiran, Boaz Barak
Abstract
We revisit and extend model stitching (Lenc & Vedaldi 2015) as a methodology to study the internal representations of neural networks. Given two trained and frozen models A and B, we consider a "stitched model" formed by connecting the bottom-layers of A to the top-layers of B, with a simple trainable layer between them. We argue that model stitching is a powerful and perhaps under-appreciated tool, which reveals aspects of representations that measures such as centered kernel alignment (CKA) cannot. Through extensive experiments, we use model stitching to obtain quantitative verifications for intuitive statements such as "good networks learn similar representations", by demonstrating that good networks of the same architecture, but trained in very different ways (e.g.: supervised vs. self-supervised learning), can be stitched to each other without drop in performance. We also give evidence for the intuition that "more is better" by showing that representations learnt with (1) more data, (2) bigger width, or (3) more training time can be "plugged in" to weaker models to improve performance. Finally, our experiments reveal a new structural property of SGD which we call "stitching connectivity", akin to mode-connectivity: typical minima reached by SGD can all be stitched to each other with minimal change in accuracy. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f1f7c80-4f54-4b0f-9b56-fba6dfba79b7Cited by top-tier papers69
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon et al.ICML 2024 · 197 citations
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye et al.NeurIPS 2022 · 162 citations
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 144 citations
- Similarity and Matching of Neural Network RepresentationsAdrián Csiszárik, Péter Korösi-Szabó, Ákos K. Matszangosz, Gergely Papp et al.NeurIPS 2021 · 105 citations
- Towards Personalized Federated Learning via Heterogeneous Model ReassemblyJiaqi Wang, Xingyi Yang, Suhan Cui, Liwei Che et al.NeurIPS 2023 · 102 citations
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
Related papers
- How Not to Stitch Representations to Measure Similarity: Task Loss Matching Versus Direct MatchingAndrás Balogh, Márk JelasityAAAI 2025 · 3 citations
- Towards a learning theory of representation alignmentFrancesco Insulla, Shuo Huang, Lorenzo RosascoICLR 2025
- Functional Alignment Can Mislead: Examining Model StitchingDamian Smith, Harvey Mannering, Antonia MarcuICML 2025
- Reliability of CKA as a Similarity Measure in Deep LearningMohammadReza Davari, Stefan Horoi, Amine Natik, Guillaume Lajoie et al.ICLR 2023 · 3 citations
- On the Functional Similarity of Robust and Non-Robust Neural RepresentationsAndrás Balogh, Márk JelasityICML 2023 · 4 citations
