Understanding the Generalization Benefit of Normalization Layers: Sharpness Reduction
Kaifeng Lyu, Zhiyuan Li, Sanjeev Arora
摘要
Normalization layers (e.g., Batch Normalization, Layer Normalization) were introduced to help with optimization difficulties in very deep nets, but they clearly also help generalization, even in not-so-deep nets. Motivated by the long-held belief that flatter minima lead to better generalization, this paper gives mathematical analysis and supporting experiments suggesting that normalization (together with accompanying weight-decay) encourages GD to reduce the sharpness of loss surface. Here "sharpness" is carefully defined given that the loss is scale-invariant, a known consequence of normalization. Specifically, for a fairly broad class of neural nets with normalization, our theory explains how GD with a finite learning rate enters the so-called Edge of Stability (EoS) regime, and characterizes the trajectory of GD in this regime via a continuous sharpness-reduction flow. w 2 2 , so WD is in effect trying to enlarge the gradient and Hessian in training. This makes the training dynamics very different from unnormalized nets and requires revisiting classical convergence analyses [77, 78, 84, 80]. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper59
- Revisiting Weighted Aggregation in Federated Learning with Neural NetworksZexi Li, Tao Lin, Xinyi Shang, Chao WuICML 2023 · 被引用 119 次
- A Modern Look at the Relationship between Sharpness and GeneralizationMaksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein 等ICML 2023 · 被引用 92 次
- Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce GrokkingKaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du 等ICLR 2024 · 被引用 71 次
- Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of StabilityZixuan Wang, Zhouzi Li, Jian LiNeurIPS 2022 · 被引用 71 次
- Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationKaiyue Wen, Zhiyuan Li, Tengyu MaNeurIPS 2023 · 被引用 53 次
它引用的顶会 Paper45
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
相关 Paper
- Fast Equilibrium of SGD in Generic SituationsZhiyuan Liu, Yi Wang, Zhiren WangICLR 2024 · 被引用 1 次
- Optimization Theory for ReLU Neural Networks Trained with Normalization LayersYonatan Dukler, Quanquan Gu, Guido MontúfarICML 2020 · 被引用 30 次
- Understanding the Disharmony between Weight Normalization Family and Weight DecayXiang Li, Shuo Chen, Jian YangAAAI 2020 · 被引用 18 次
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou 等ICLR 2023 · 被引用 1 次
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt 等ICML 2026
