Should Under-parameterized Student Networks Copy or Average Teacher Weights?
Berfin Simsek, Amire Bendjeddou, Wulfram Gerstner, Johanni Brea
Abstract
Any continuous function can be approximated arbitrarily well by a neural network with sufficiently many neurons . We consider the case when itself is a neural network with one hidden layer and neurons. Approximating with a neural network with neurons can thus be seen as fitting an under-parameterized"student"network with neurons to a"teacher"network with neurons. As the student has fewer neurons than the teacher, it is unclear, whether each of the student neurons should copy one of the teacher neurons or rather average a group of teacher neurons. For shallow neural networks with erf activation function and for the standard Gaussian input distribution, we prove that"copy-average"configurations are critical points if the teacher's incoming vectors are orthonormal and its outgoing weights are unitary. Moreover, the optimum among such configurations is reached when student neurons each copy one teacher neuron and the -th student neuron averages the remaining teacher neurons. For the student network with neuron, we provide additionally a closed-form solution of the non-trivial critical point(s) for commonly used activation functions through solving an equivalent constrained optimization problem. Empirically, we find for the erf activation function that gradient flow converges either to the optimal copy-average critical point or to another point where each student neuron approximately copies a different teacher neuron. Finally, we find similar results for the ReLU activation function, suggesting that the optimal solution of underparameterized networks has a universal structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db5df8da-d56c-4370-aca1-7f6624f0846cCited by top-tier papers3
- Expand-and-Cluster: Parameter Recovery of Neural NetworksFlavio Martinelli, Berfin Simsek, Wulfram Gerstner, Johanni BreaICML 2024 · 15 citations
- Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural NetworksBinghui Li, Zhixuan Pan, Kaifeng Lyu, Jian LiICLR 2025
- How Gradient descent balances features: A dynamical analysis for two-layer neural networksZhenyu Zhu, Fanghui Liu, Volkan CevherICLR 2025
Builds on6
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 94 citations
- Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index ModelsAlex Damian, Eshaan Nichani, Rong Ge, Jason D. LeeNeurIPS 2023 · 67 citations
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- Learning a Single Neuron with Bias Using Gradient DescentGal Vardi, Gilad Yehudai, Ohad ShamirNeurIPS 2021 · 23 citations
Related papers
- ReLU Network with Width d+O(1) Can Achieve Optimal Approximation RateChenghao Liu, Minghua ChenICML 2024 · 3 citations
- Universal approximation power of deep residual neural networks via nonlinear control theoryPaulo Tabuada, Bahman GharesifardICLR 2021 · 31 citations
- Achieve the Minimum Width of Neural Networks for Universal ApproximationYongqiang CaiICLR 2023 · 4 citations
- Non-Vacuous Generalisation Bounds for Shallow Neural NetworksFelix Biggs, Benjamin GuedjICML 2022 · 29 citations
- SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal StructuresJulian Kranz, Davide Gallon, Steffen Dereich, Arnulf JentzenNeurIPS 2025 · 5 citations
