On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
Zeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato, Masashi Sugiyama
Abstract
Weight decay is a simple yet powerful regularization technique that has been very widely used in training of deep neural networks (DNNs). While weight decay has attracted much attention, previous studies fail to discover some overlooked pitfalls on large gradient norms resulted by weight decay. In this paper, we discover that, weight decay can unfortunately lead to large gradient norms at the final phase (or the terminated solution) of training, which often indicates bad convergence and poor generalization. To mitigate the gradient-norm-centered pitfalls, we present the first practical scheduler for weight decay, called the Scheduled Weight Decay (SWD) method that can dynamically adjust the weight decay strength according to the gradient norm and significantly penalize large gradient norms during training. Our experiments also support that SWD indeed mitigates large gradient norms and often significantly outperforms the conventional constant weight decay strategy for Adaptive Moment Estimation (Adam).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e3fe096-fdb5-41f9-b1dd-d97fffb72c96Cited by top-tier papers16
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens et al.NeurIPS 2024 · 69 citations
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 41 citations
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
- QVGen: Pushing the Limit of Quantized Video Generative ModelsYushi Huang, Ruihao Gong, Jing Liu, Yifu Ding et al.ICLR 2026 · 22 citations
- On the Overlooked Structure of Stochastic GradientsZeke Xie, Qian-Yuan Tang, Mingming Sun, Ping LiNeurIPS 2023 · 18 citations
Builds on16
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
Related papers
- Understanding Decoupled and Early Weight DecayJohan Bjorck, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 37 citations
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
- Investigating the Role of Weight Decay in Enhancing Nonconvex SGDTao Sun, Yuhao Huang, Li Shen, Kele Xu et al.CVPR 2025
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Understanding the Generalization of Adam in Learning Neural Networks with Proper RegularizationDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICLR 2023 · 6 citations
