Towards Understanding the Condensation of Neural Networks at Initial Training
Hanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang, Zhi-Qin John Xu
Abstract
Empirical works show that for ReLU neural networks (NNs) with small initialization, input weights of hidden neurons (the input weight of a hidden neuron consists of the weight from its input layer to the hidden neuron and its bias term) condense onto isolated orientations. The condensation dynamics implies that the training implicitly regularizes a NN towards one with much smaller effective size. In this work, we illustrate the formation of the condensation in multi-layer fully connected NNs and show that the maximal number of condensed orientations in the initial training stage is twice the multiplicity of the activation function, where "multiplicity" indicates the multiple roots of activation function at origin. Our theoretical analysis confirms experiments for two cases, one is for the activation function of multiplicity one with arbitrary dimension input, which contains many common activation functions, and the other is for the layer with one-dimensional input and arbitrary multiplicity. This work makes a step towards understanding how small initialization leads NNs to condensation at the initial training stage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a117748-d6a6-4c8b-b5ca-80a239926af3Cited by top-tier papers8
- Gradient Alignment in Physics-informed Neural Networks: A Second-Order Optimization PerspectiveSifan Wang, Ananyae Kumar Bhartari, Bowen Li, Paris PerdikarisNeurIPS 2025 · 100 citations
- Understanding Multi-phase Optimization Dynamics and Rich Nonlinear Behaviors of ReLU NetworksMingze Wang, Chao MaNeurIPS 2023 · 17 citations
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training DynamicsZheng-An Chen, Tao LuoNeurIPS 2025 · 13 citations
- Simplicity Bias of Two-Layer Networks beyond Linearly Separable DataNikita Tsoy, Nikola KonstantinovICML 2024 · 12 citations
- Connectivity Shapes Implicit Regularization in Matrix Factorization Models for Matrix CompletionZhiwei Bai, Jiajie Zhao, Yaoyu ZhangNeurIPS 2024 · 6 citations
Builds on11
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro et al.ICML 2021 · 136 citations
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 94 citations
Related papers
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 92 citations
- Empirical Phase Diagram for Three-layer Neural Networks with Infinite WidthHanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo et al.NeurIPS 2022 · 22 citations
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 15 citations
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 35 citations
- Early Neuron Alignment in Two-layer ReLU Networks with Small InitializationHancheng Min, Enrique Mallada, René VidalICLR 2024 · 31 citations
