Functional Alignment Can Mislead: Examining Model Stitching
Damian Smith, Harvey Mannering, Antonia Marcu
Abstract
A common belief in the representational comparison literature is that if two representations can be functionally aligned, they must capture similar information. In this paper we focus on model stitching and show that models can be functionally aligned, but represent very different information. Firstly, we show that discriminative models with very different biases can be stitched together. We then show that models trained to solve entirely different tasks on different data modalities, and even representations in the form of clustered random noise, can be successfully stitched into MNIST or ImageNet-trained models. We proceed by showing that alignments can also be found in the case of autoencoders where the encoder and decoder are trained on different tasks. We end with a discussion of the wider impact of our results on the community's current beliefs. Overall, our paper draws attention to the need to correctly interpret the results of such functional similarity measures and highlights the need for approaches that capture informational similarity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b658e29f-822c-457c-b00d-6f4ad920c177Cited by top-tier papers2
- Revisiting Model Stitching In the Foundation Model EraZheda Mai, Ke Zhang, Fu-En Wang, Zixiao Ken Wang et al.CVPR 2026 · 2 citations
- Grounding Functional Similarity by Invariance-Aware Model StitchingIoannis Athanasiadis, Anmar Karmush, Michael FelsbergICML 2026 · 1 citation
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Learning De-biased Representations with Biased RepresentationsHyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo et al.ICML 2020 · 332 citations
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 323 citations
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 253 citations
Related papers
- On the Functional Similarity of Robust and Non-Robust Neural RepresentationsAndrás Balogh, Márk JelasityICML 2023 · 4 citations
- Connecting Neural Models Latent Geometries with Relative Geodesic RepresentationsHanlin Yu, Berfin Inal, Georgios Arvanitidis, Søren Hauberg et al.NeurIPS 2025 · 5 citations
- Latent Space Translation via Semantic AlignmentValentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero et al.NeurIPS 2023 · 59 citations
- Similarity and Matching of Neural Network RepresentationsAdrián Csiszárik, Péter Korösi-Szabó, Ákos K. Matszangosz, Gergely Papp et al.NeurIPS 2021 · 105 citations
- How Not to Stitch Representations to Measure Similarity: Task Loss Matching Versus Direct MatchingAndrás Balogh, Márk JelasityAAAI 2025 · 3 citations
