Expand-and-Cluster: Parameter Recovery of Neural Networks
Flavio Martinelli, Berfin Simsek, Wulfram Gerstner, Johanni Brea
Abstract
Can we identify the weights of a neural network by probing its input-output mapping? At first glance, this problem seems to have many solutions because of permutation, overparameterisation and activation function symmetries. Yet, we show that the incoming weight vector of each neuron is identifiable up to sign or scaling, depending on the activation function. Our novel method 'Expand-and-Cluster' can identify layer sizes and weights of a target network for all commonly used activation functions. Expand-and-Cluster consists of two phases: (i) to relax the non-convex optimisation problem, we train multiple overparameterised student networks to best imitate the target function; (ii) to reverse engineer the target network's weights, we employ an ad-hoc clustering procedure that reveals the learnt weight vectors shared between students -- these correspond to the target weight vectors. We demonstrate successful weights and size recovery of trained shallow and deep networks with less than 10% overhead in the layer size and describe an `ease-of-identifiability' axis by analysing 150 synthetic problems of variable difficulty.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60514c3b-57ea-4cc9-a0af-baf3fe36290eCited by top-tier papers9
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural NetworksAnn Huang, Satpreet Harcharan Singh, Flavio Martinelli, Kanaka RajanNeurIPS 2025 · 22 citations
- Identifiability of Deep Polynomial Neural NetworksKonstantin Usevich, Ricardo Augusto Borsoi, Clara Dérand, Marianne ClauselNeurIPS 2025 · 21 citations
- Beyond Slow Signs in High-fidelity Model ExtractionHanna Foerster, Robert Mullins, Ilia Shumailov, Jamie HayesNeurIPS 2024 · 19 citations
- Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label SettingNicholas Carlini, Jorge Chávez-Saab, Anna Hambitzer, Francisco Rodríguez-Henríquez et al.EUROCRYPT 2025 · 10 citations
- Flat Channels to Infinity in Neural Loss LandscapesFlavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner et al.NeurIPS 2025 · 6 citations
Builds on22
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
Related papers
- Should Under-parameterized Student Networks Copy or Average Teacher Weights?Berfin Simsek, Amire Bendjeddou, Wulfram Gerstner, Johanni BreaNeurIPS 2023 · 14 citations
- Reverse-engineering deep ReLU networksDavid Rolnick, Konrad P. KordingICML 2020 · 121 citations
- Identifiable Equivariant Networks are Layerwise EquivariantVahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman, Kathlén KohnICML 2026
- Functional vs. parametric equivalence of ReLU networksMary Phuong, Christoph H. LampertICLR 2020 · 53 citations
- Span Recovery for Deep Neural Networks with Applications to Input ObfuscationRajesh Jayaram, David P. Woodruff, Qiuyi ZhangICLR 2020 · 6 citations
