The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks
Lénaïc Chizat, Praneeth Netrapalli
摘要
Deep learning succeeds by doing hierarchical feature learning, yet tuning hyper-parameters (HP) such as initialization scales, learning rates etc., only give indirect control over this behavior. In this paper, we introduce a key notion to predict and control feature learning: the angle between the feature updates and the backward pass (at layer index ). We show that the magnitude of feature updates after one GD step, at any training time, can be expressed via a simple and general feature speed formula in terms of this angle , the loss decay, and the magnitude of the backward pass. This angle is controlled by the conditioning of the layer-to-layer Jacobians and at random initialization, it is determined by the spectrum of a certain kernel, which coincides with the Neural Tangent Kernel when . Given , the feature speed formula provides us with rules to adjust HPs (scales and learning rates) so as to satisfy certain dynamical properties, such as feature learning and loss decay. We investigate the implications of our approach for ReLU MLPs and ResNets in the large width-then-depth limit. Relying on prior work, we show that in ReLU MLPs with iid initialization, the angle degenerates with depth as . In contrast, ResNets with branch scale maintain a non-degenerate angle . We use these insights to recover key properties of known HP scalings and also to introduce a new HP scaling for large depth ReLU MLPs with favorable theoretical properties.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 被引用 21 次
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and TimeBlake Bordelon, Mary I. Letey, Cengiz PehlevanICLR 2026 · 被引用 14 次
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 被引用 8 次
- On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingMoritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena Chennuru VankadaraNeurIPS 2025 · 被引用 7 次
- Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural NetworksHaosong Zhang, Shenxi Wu, Xingjian Ma, Shirui Bian 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper11
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 被引用 110 次
- A mathematical model for automatic differentiation in machine learningJérôme Bolte, Edouard PauwelsNeurIPS 2020 · 被引用 84 次
- Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling LimitBlake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin 等ICLR 2024 · 被引用 54 次
相关 Paper
- Super Consistency of Neural Network Landscapes and Learning Rate TransferLorenzo Noci, Alexandru Meterez, Thomas Hofmann, Antonio OrvietoNeurIPS 2024 · 被引用 25 次
- Maximal Initial Learning Rates in Deep ReLU NetworksGaurav Iyer, Boris Hanin, David RolnickICML 2023 · 被引用 14 次
- Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in GeneralizationTaesun Yeom, Taehyeok Ha, Jaeho LeeICML 2026
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- On the Random Conjugate Kernel and Neural Tangent KernelZhengmian Hu, Heng HuangICML 2021 · 被引用 15 次
