Understanding Multi-phase Optimization Dynamics and Rich Nonlinear Behaviors of ReLU Networks
Mingze Wang, Chao Ma
摘要
The training process of ReLU neural networks often exhibits complicated nonlinear phenomena. The nonlinearity of models and non-convexity of loss pose significant challenges for theoretical analysis. Therefore, most previous theoretical works on the optimization dynamics of neural networks focus either on local analysis (like the end of training) or approximate linear models (like Neural Tangent Kernel). In this work, we conduct a complete theoretical characterization of the training process of a two-layer ReLU network trained by Gradient Flow on a linearly separable data. In this specific setting, our analysis captures the whole optimization process starting from random initialization to final convergence. Despite the relatively simple model and data that we studied, we reveal four different phases from the whole training process showing a general simplifying-to-complicating learning trend. Specific nonlinear behaviors can also be precisely identified and captured theoretically, such as initial condensation, saddle-to-plateau dynamics, plateau escape, changes of activation patterns, learning with increasing complexity, etc. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen 等NeurIPS 2024 · 被引用 48 次
- Understanding the Expressive Power and Mechanisms of Transformer for Sequence ModelingMingze Wang, Weinan ENeurIPS 2024 · 被引用 32 次
- Early Neuron Alignment in Two-layer ReLU Networks with Small InitializationHancheng Min, Enrique Mallada, René VidalICLR 2024 · 被引用 31 次
- Improving Generalization and Convergence by Enhancing Implicit RegularizationMingze Wang, Jinbo Wang, Haotian He, Zilin Wang 等NeurIPS 2024 · 被引用 21 次
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 被引用 15 次
它引用的顶会 Paper28
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 等NeurIPS 2021 · 被引用 303 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 被引用 182 次
相关 Paper
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 被引用 92 次
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable DataHancheng Min, Zhihui Zhu, René VidalNeurIPS 2025 · 被引用 3 次
- Empirical Phase Diagram for Three-layer Neural Networks with Infinite WidthHanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo 等NeurIPS 2022 · 被引用 22 次
- Towards Understanding the Condensation of Neural Networks at Initial TrainingHanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang 等NeurIPS 2022 · 被引用 42 次
- Convergence of the Gradient Flow for Shallow ReLU Networks on Weakly Interacting DataLéo Dana, Loucas Pillaud-Vivien, Francis BachNeurIPS 2025 · 被引用 1 次
