On the training dynamics of deep networks with regularization
Aitor Lewkowycz, Guy Gur-Ari
Abstract
We study the role of regularization in deep learning, and uncover simple relations between the performance of the model, the coefficient, the learning rate, and the number of training steps. These empirical relations hold when the network is overparameterized. They can be used to predict the optimal regularization parameter of a given model. In addition, based on these observations we propose a dynamical schedule for the regularization parameter that improves performance and speeds up training. We test these proposals in modern image classification settings. Finally, we show that these empirical relations can be understood theoretically in the context of infinitely wide networks. We derive the gradient flow dynamics of such networks, and compare the role of regularization in this context with that of linear models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf5d1d27-438c-4834-810d-da68b575c8ecCited by top-tier papers24
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
- Revisiting Weighted Aggregation in Federated Learning with Neural NetworksZexi Li, Tao Lin, Xinyi Shang, Chao WuICML 2023 · 119 citations
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
Builds on3
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 127 citations
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 93 citations
Related papers
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning RateZhiqi Bu, Shiyun Xu, Jialin MaoICLR 2026 · 4 citations
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 77 citations
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato et al.NeurIPS 2023 · 38 citations
- The Implicit Regularization of Momentum Gradient Descent in Overparametrized ModelsLi Wang, Zhiguo Fu, Yingcong Zhou, Zili YanAAAI 2023 · 9 citations
- Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter TransferBlake Bordelon, Cengiz PehlevanICML 2025
