Reverse-engineering deep ReLU networks
David Rolnick, Konrad P. Kording
Abstract
It has been widely assumed that a neural network cannot be recovered from its outputs, as the network depends on its parameters in a highly nonlinear way. Here, we prove that in fact it is often possible to identify the architecture, weights, and biases of an unknown deep ReLU network by observing only its output. Every ReLU network defines a piecewise linear function, where the boundaries between linear regions correspond to inputs for which some neuron in the network switches between inactive and active ReLU states. By dissecting the set of region boundaries into components associated with particular neurons, we show both theoretically and empirically that it is possible to recover the weights of neurons and their arrangement within the network, up to isomorphism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d2b694d-0ea4-4ffc-a4e3-821d211f3a11Cited by top-tier papers42
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir et al.NeurIPS 2022 · 196 citations
- DeepSteal: Advanced Model Extractions Leveraging Efficient Weight Stealing in MemoriesAdnan Siraj Rakin, Md Hafizul Islam Chowdhuryy, Fan Yao, Deliang FanS&P 2022 · 163 citations
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke et al.ICML 2024 · 157 citations
- Muse: Secure Inference Resilient to Malicious ClientsRyan Lehmkuhl, Pratyush Mishra, Akshayaram Srinivasan, Raluca Ada PopaUSENIX Security 2021 · 115 citations
- Cryptanalytic Extraction of Neural Network ModelsNicholas Carlini, Matthew Jagielski, Ilya MironovCRYPTO 2020 · 109 citations
Builds on1
Related papers
- Polynomial Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label SettingNicholas Carlini, Jorge Chávez-Saab, Anna Hambitzer, Francisco Rodríguez-Henríquez et al.EUROCRYPT 2025 · 10 citations
- Cryptanalytic Extraction of Deep Neural Networks with Non-linear ActivationsRoderick Asselineau, Patrick Derbez, Pierre-Alain Fouque, Brice MinaudCRYPTO 2026 · 8 citations
- Efficient Algorithms for Learning Depth-2 Neural Networks with General ReLU ActivationsPranjal Awasthi, Alex Tang, Aravindan VijayaraghavanNeurIPS 2021 · 24 citations
- The phase diagram of approximation rates for deep neural networksDmitry Yarotsky, Anton ZhevnerchukNeurIPS 2020 · 156 citations
- Compelling ReLU Networks to Exhibit Exponentially Many Linear Regions at Initialization and During TrainingMax Milkert, David Hyde, Forrest J. LaineICML 2025
