Expand-and-Cluster: Parameter Recovery of Neural Networks
Flavio Martinelli, Berfin Simsek, Wulfram Gerstner, Johanni Brea
摘要
Can we identify the weights of a neural network by probing its input-output mapping? At first glance, this problem seems to have many solutions because of permutation, overparameterisation and activation function symmetries. Yet, we show that the incoming weight vector of each neuron is identifiable up to sign or scaling, depending on the activation function. Our novel method 'Expand-and-Cluster' can identify layer sizes and weights of a target network for all commonly used activation functions. Expand-and-Cluster consists of two phases: (i) to relax the non-convex optimisation problem, we train multiple overparameterised student networks to best imitate the target function; (ii) to reverse engineer the target network's weights, we employ an ad-hoc clustering procedure that reveals the learnt weight vectors shared between students -- these correspond to the target weight vectors. We demonstrate successful weights and size recovery of trained shallow and deep networks with less than 10% overhead in the layer size and describe an `ease-of-identifiability' axis by analysing 150 synthetic problems of variable difficulty.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural NetworksAnn Huang, Satpreet Harcharan Singh, Flavio Martinelli, Kanaka RajanNeurIPS 2025 · 被引用 22 次
- Identifiability of Deep Polynomial Neural NetworksKonstantin Usevich, Ricardo Augusto Borsoi, Clara Dérand, Marianne ClauselNeurIPS 2025 · 被引用 21 次
- Beyond Slow Signs in High-fidelity Model ExtractionHanna Foerster, Robert Mullins, Ilia Shumailov, Jamie HayesNeurIPS 2024 · 被引用 19 次
- Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label SettingNicholas Carlini, Jorge Chávez-Saab, Anna Hambitzer, Francisco Rodríguez-Henríquez 等EUROCRYPT 2025 · 被引用 10 次
- Flat Channels to Infinity in Neural Loss LandscapesFlavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper22
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos 等ICLR 2020 · 被引用 1,368 次
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 被引用 330 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 被引用 301 次
相关 Paper
- Should Under-parameterized Student Networks Copy or Average Teacher Weights?Berfin Simsek, Amire Bendjeddou, Wulfram Gerstner, Johanni BreaNeurIPS 2023 · 被引用 14 次
- Reverse-engineering deep ReLU networksDavid Rolnick, Konrad P. KordingICML 2020 · 被引用 121 次
- Identifiable Equivariant Networks are Layerwise EquivariantVahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman, Kathlén KohnICML 2026
- Functional vs. parametric equivalence of ReLU networksMary Phuong, Christoph H. LampertICLR 2020 · 被引用 53 次
- Span Recovery for Deep Neural Networks with Applications to Input ObfuscationRajesh Jayaram, David P. Woodruff, Qiuyi ZhangICLR 2020 · 被引用 6 次
