On the Functional Similarity of Robust and Non-Robust Neural Representations
András Balogh, Márk Jelasity
Abstract
Model stitching-where the internal representations of two neural networks are aligned linearlyhelped demonstrate that the representations of different neural networks for the same task are surprisingly similar in a functional sense. At the same time, the representations of adversarially robust networks are considered to be different from non-robust representations. For example, robust image classifiers are invertible, while nonrobust networks are not. Here, we investigate the functional similarity of robust and non-robust representations for image classification with the help of model stitching. We find that robust and non-robust networks indeed have different representations. However, these representations are compatible regarding accuracy. From the point of view of robust accuracy, compatibility decreases quickly after the first few layers but the representations become compatible again in the last layers, in the sense that the properties of the front model can be recovered. Moreover, this is true even in the case of cross-task stitching. Our results suggest that stitching in the initial, preprocessing layers and the final, abstract layers test different kinds of compatibilities. In particular, the final layers are easy to match, because their representations depend mostly on the same abstract task specification, in our case, the classification of the input into n classes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- How Not to Stitch Representations to Measure Similarity: Task Loss Matching Versus Direct MatchingAndrás Balogh, Márk JelasityAAAI 2025 · 3 citations
- Grounding Functional Similarity by Invariance-Aware Model StitchingIoannis Athanasiadis, Anmar Karmush, Michael FelsbergICML 2026 · 1 citation
- Less Is More: Rethinking Parameter-Efficient Fine-Tuning from a Subtractive PerspectiveTianqi Jiang, Liu Yang, Xi-Le Zhao, Zixuan Qin et al.AAAI 2026
Builds on10
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 253 citations
- Linear Mode Connectivity in Multitask and Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Dilan Görür, Razvan Pascanu et al.ICLR 2021 · 176 citations
- Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai et al.ICLR 2022 · 150 citations
Related papers
- Functional Alignment Can Mislead: Examining Model StitchingDamian Smith, Harvey Mannering, Antonia MarcuICML 2025
- Similarity and Matching of Neural Network RepresentationsAdrián Csiszárik, Péter Korösi-Szabó, Ákos K. Matszangosz, Gergely Papp et al.NeurIPS 2021 · 105 citations
- Adversarial Training Reduces Information and Improves TransferabilityMatteo Terzi, Alessandro Achille, Marco Maggipinto, Gian Antonio SustoAAAI 2021 · 25 citations
- Neural Representations Reveal Distinct Modes of Class Fitting in Residual Convolutional NetworksMichal Jamroz, Marcin KurdzielAAAI 2023
- From Bricks to Bridges: Product of Invariances to Enhance Latent Space CommunicationIrene Cannistraci, Luca Moschella, Marco Fumero, Valentino Maiorca et al.ICLR 2024 · 22 citations
