Empirical Phase Diagram for Three-layer Neural Networks with Infinite Width
Hanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo, Yaoyu Zhang, Zhi-Qin John Xu
Abstract
Substantial work indicates that the dynamics of neural networks (NNs) is closely related to their initialization of parameters. Inspired by the phase diagram for two-layer ReLU NNs with infinite width (Luo et al., 2021), we make a step towards drawing a phase diagram for three-layer ReLU NNs with infinite width. First, we derive a normalized gradient flow for three-layer ReLU NNs and obtain two key independent quantities to distinguish different dynamical regimes for common initialization methods. With carefully designed experiments and a large computation cost, for both synthetic datasets and real datasets, we find that the dynamics of each layer also could be divided into a linear regime and a condensed regime, separated by a critical regime. The criteria is the relative change of input weights (the input weight of a hidden neuron consists of the weight from its input layer to the hidden neuron and its bias term) as the width approaches infinity during the training, which tends to , and , respectively. In addition, we also demonstrate that different layers can lie in different dynamical regimes in a training process within a deep NN. In the condensed regime, we also observe the condensation of weights in isolated orientations with low complexity. Through experiments under three-layer condition, our phase diagram suggests a complicated dynamical regimes consisting of three possible regimes, together with their mixture, for deep NNs and provides a guidance for studying deep NNs in different initialization regimes, which reveals the possibility of completely different dynamics emerging within a deep NN for its different layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Towards Understanding the Condensation of Neural Networks at Initial TrainingHanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang et al.NeurIPS 2022 · 42 citations
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 21 citations
- Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or MemorizingZhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang et al.NeurIPS 2024 · 20 citations
- Understanding Multi-phase Optimization Dynamics and Rich Nonlinear Behaviors of ReLU NetworksMingze Wang, Chao MaNeurIPS 2023 · 17 citations
- Multi-Layer Neural Networks as Trainable Ladders of Hilbert SpacesZhengdao ChenICML 2023 · 4 citations
Builds on5
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 94 citations
- Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training DynamicsGreg Yang, Etai LittwinICML 2021 · 81 citations
- Deep Frequency Principle Towards Understanding Why Deeper Learning Is FasterZhiqin John Xu, Hanxu ZhouAAAI 2021 · 67 citations
- Embedding Principle of Loss Landscape of Deep Neural NetworksYaoyu Zhang, Zhongwang Zhang, Tao Luo, Zhi-Qin John XuNeurIPS 2021 · 48 citations
- An analytic theory of shallow networks dynamics for hinge loss classificationFranco Pellegrini, Giulio BiroliNeurIPS 2020 · 19 citations
Related papers
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth et al.ICML 2021 · 85 citations
- Mixed Dynamics In Linear Networks: Unifying the Lazy and Active RegimesZhenfeng Tu, Santiago Aranguri, Arthur JacotNeurIPS 2024 · 18 citations
- Three Mechanisms of Feature Learning in a Linear NetworkYizhou Xu, Ziyin LiuICLR 2025
- From Lazy to Rich: Exact Learning Dynamics in Deep Linear NetworksClémentine Carla Juliette Dominé, Nicolas Anguita, Alexandra Maria Proca, Lukas Braun et al.ICLR 2025
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 92 citations
