Understanding the Disharmony between Weight Normalization Family and Weight Decay
Xiang Li, Shuo Chen, Jian Yang
摘要
The merits of fast convergence and potentially better performance of the weight normalization family have drawn increasing attention in recent years. These methods use standardization or normalization that changes the weight W to W , which makes W independent to the magnitude of W . Surprisingly, W must be decayed during gradient descent, otherwise we will observe a severe under-fitting problem, which is very counter-intuitive since weight decay is widely known to prevent deep networks from over-fitting. Moreover, if we substitute (e.g., weight normalization) it is observed that the regularization term 1 2 λ||W || 2 will be canceled as a constant 1 2 λ in the optimization objective. Therefore, to decay W , we need to explicitly append this term: 1 2 λ||W || 2 . In this paper, we theoretically prove that 1 2 λ||W || 2 merely modulates the effective learning rate for improving objective optimization, and has no influence on generalization when the weight normalization family is compositely employed. Furthermore, we also expose several critical problems when introducing weight decay term to weight normalization family, including the missing of global minimum and training instability. To address these problems, we propose an -shifted L 2 regularizer, which shifts the L 2 objective by a positive constant . Such a simple operation can theoretically guarantee the existence of global minimum, while preventing the network weights from being too small and thus avoiding gradient float overflow. It significantly improves the training stability and can achieve slightly better performance in our practice. The effectiveness of -shifted L 2 regularizer is comprehensively validated on the ImageNet, CIFAR-100, and COCO datasets. Our codes and pretrained models will be released in https://github.com/implus/PytorchInsight .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Revisiting Weighted Aggregation in Federated Learning with Neural NetworksZexi Li, Tao Lin, Xinyi Shang, Chao WuICML 2023 · 被引用 119 次
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens 等NeurIPS 2024 · 被引用 69 次
- On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight DecayEkaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin 等NeurIPS 2021 · 被引用 30 次
- Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three RegimesMaxim Kodryan, Ekaterina Lobacheva, Maksim Nakhodnov, Dmitry P. VetrovNeurIPS 2022 · 被引用 25 次
- Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality RegularizationJunlin He, Jinxiao Du, Wei MaNeurIPS 2024 · 被引用 19 次
它引用的顶会 Paper1
相关 Paper
- Understanding Decoupled and Early Weight DecayJohan Bjorck, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 被引用 37 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
