Analysis of one-hidden-layer neural networks via the resolvent method
Vanessa Piccolo, Dominik Schröder
Abstract
In this work, we investigate the asymptotic spectral density of the random feature matrix M = Y Y * with Y = f (W X) generated by a single-hidden-layer neural network, where W and X are random rectangular matrices with i.i.d. centred entries and f is a non-linear smooth function which is applied entry-wise. We prove that the Stieltjes transform of the limiting spectral distribution approximately satisfies a quartic self-consistent equation, which is exactly the equation obtained by Pennington and Worah [22] and Benigni and Péché [6] with the moment method. We extend the previous results to the case of additive bias Y = f (W X + B) with B being an independent rank-one Gaussian random matrix, closer modelling the neural network infrastructures encountered in practice. Our key finding is that in the case of additive bias it is impossible to choose an activation function preserving the layer-to-layer singular value distribution, in sharp contrast to the bias-free case where a simple integral constraint is sufficient to achieve isospectrality. To obtain the asymptotics for the empirical spectral density we follow the resolvent method from random matrix theory via the cumulant expansion. We find that this approach is more robust and less combinatorial than the moment method and expect that it will apply also for models where the combinatorics of the former become intractable. The resolvent method has been widely employed, but compared to previous works, it is applied here to non-linear random matrices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59a80baf-4ac4-40f8-b02a-824d8292063bCited by top-tier papers2
- Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterizationSimone Bombari, Mohammad Hossein Amani, Marco MondelliNeurIPS 2022 · 45 citations
- Low Rank Gradients and Where to Find ThemRishi Sonthalia, Michael Murray, Guido F. MontúfarNeurIPS 2025 · 5 citations
Builds on1
Related papers
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 51 citations
- Deep Equilibrium Models are Almost Equivalent to Not-so-deep Explicit Models for High-dimensional Gaussian MixturesZenan Ling, Longbo Li, Zhanbo Feng, Yixuan Zhang et al.ICML 2024 · 6 citations
- Eigen Analysis of Conjugate Kernel and Neural Tangent KernelXiangchao Li, Xiao Han, Qing YangICML 2025
- Characterizing the spectrum of the NTK via a power series expansionMichael Murray, Hui Jin, Benjamin Bowman, Guido MontúfarICLR 2023 · 2 citations
- Random matrices in service of ML footprint: ternary random features with no performance lossHafiz Tiomoko Ali, Zhenyu Liao, Romain CouilletICLR 2022 · 8 citations
