Expected Gradients of Maxout Networks and Consequences to Parameter Initialization
Hanna Tseran, Guido Montúfar
摘要
We study the gradients of a maxout network with respect to inputs and parameters and obtain bounds for the moments depending on the architecture and the parameter distribution. We observe that the distribution of the input-output Jacobian depends on the input, which complicates a stable parameter initialization. Based on the moments of the gradients, we formulate parameter initialization strategies that avoid vanishing and exploding gradients in wide networks. Experiments with deep fully-connected and convolutional networks show that this strategy improves SGD and Adam training of deep maxout networks. In addition, we obtain refined bounds on the expected number of linear regions, results on the expected curve length distortion, and results on the NTK. 3
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- On the Expected Complexity of Maxout NetworksHanna Tseran, Guido MontúfarNeurIPS 2021 · 被引用 19 次
- Better Training using Weight-Constrained Stochastic DynamicsBenedict J. Leimkuhler, Tiffany J. Vlaar, Timothée Pouchon, Amos J. StorkeyICML 2021 · 被引用 11 次
- Neural Tangent Kernel Beyond the Infinite-Width Limit: Effects of Depth and InitializationMariia Seleznova, Gitta KutyniokICML 2022 · 被引用 34 次
- Information Geometry of Orthogonal Initializations and TrainingPiotr Aleksander Sokól, Il Memming ParkICLR 2020 · 被引用 17 次
- Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning SystemsOri Shem-Ur, Khen Cohen, Yaron OzICLR 2026
