Deconstructing the Goldilocks Zone of Neural Network Initialization
Artem Vysogorets, Anna Dawid, Julia Kempe
摘要
The second-order properties of the training loss have a massive impact on the optimization dynamics of deep learning models. Fort & Scherlis (2019) discovered that a large excess of positive curvature and local convexity of the loss Hessian is associated with highly trainable initial points located in a region coined the "Goldilocks zone". Only a handful of subsequent studies touched upon this relationship, so it remains largely unexplained. In this paper, we present a rigorous and comprehensive analysis of the Goldilocks zone for homogeneous neural networks. In particular, we derive the fundamental condition resulting in excess of positive curvature of the loss, explaining and refining its conventionally accepted connection to the initialization norm. Further, we relate the excess of positive curvature to model confidence, low initial loss, and a previously unknown type of vanishing cross-entropy loss gradient. To understand the importance of excessive positive curvature for trainability of deep networks, we optimize fully-connected and convolutional architectures outside the Goldilocks zone and analyze the emergent behaviors. We find that strong model performance is not perfectly aligned with the Goldilocks zone, calling for further research into this relationship.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng 等ICML 2022 · 被引用 386 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- Bad Global Minima Exist and SGD Can Reach ThemShengchao Liu, Dimitris S. Papailiopoulos, Dimitris AchlioptasNeurIPS 2020 · 被引用 89 次
相关 Paper
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 被引用 15 次
- Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU NetworksArsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin GolestaniICLR 2020 · 被引用 12 次
- A Loss Curvature Perspective on Training Instabilities of Deep Learning ModelsJustin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta 等ICLR 2022 · 被引用 49 次
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 被引用 21 次
- Entropic Confinement and Mode Connectivity in Overparameterized Neural NetworksLuca di Carlo, Chase Goddard, David J. SchwabICLR 2026
