Penalizing Gradient Norm for Efficiently Improving Generalization in Deep Learning
Yang Zhao, Hao Zhang, Xiuyuan Hu
Abstract
How to train deep neural networks (DNNs) to generalize well is a central concern in deep learning, especially for severely overparameterized networks nowadays. In this paper, we propose an effective method to improve the model generalization by additionally penalizing the gradient norm of loss function during optimization. We demonstrate that confining the gradient norm of loss function could help lead the optimizers towards finding flat minima. We leverage the first-order approximation to efficiently implement the corresponding gradient to fit well in the gradient descent framework. In our experiments, we confirm that when using our methods, generalization performance of various models could be improved on different datasets. Also, we show that the recent sharpness-aware minimization method (Foret et al., 2021) is a special, but not the best, case of our method, where the best case of our method could give new state-of-art performance on these tasks. Code is available at https://github.com/zhaoyang-0204/gnp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers76
- Boosting Adversarial Transferability by Achieving Flat Local MaximaZhijin Ge, Xiaosen Wang, Hongying Liu, Fanhua Shang et al.NeurIPS 2023 · 112 citations
- Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation ApproachPeng Mi, Li Shen, Tianhe Ren, Yiyi Zhou et al.NeurIPS 2022 · 102 citations
- Improving the Model Consistency of Decentralized Federated LearningYifan Shi, Li Shen, Kang Wei, Yan Sun et al.ICML 2023 · 89 citations
- Dynamic Regularized Sharpness Aware Minimization in Federated Learning: Approaching Global Consistency and Smooth LandscapeYan Sun, Li Shen, Shixiang Chen, Liang Ding et al.ICML 2023 · 69 citations
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han et al.NeurIPS 2023 · 50 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 165 citations
- Regularizing Neural Networks via Adversarial Model PerturbationYaowei Zheng, Richong Zhang, Yongyi MaoCVPR 2021
Related papers
- Sharpness-Aware Training for FreeJiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan et al.NeurIPS 2022 · 132 citations
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao et al.KDD 2023 · 7 citations
- Differentially Private Sharpness-Aware TrainingJinseong Park, Hoki Kim, Yujin Choi, Jaewook LeeICML 2023 · 15 citations
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou et al.CVPR 2023
- Fundamental Convergence Analysis of Sharpness-Aware MinimizationPham Duy Khanh, Hoang-Chau Luong, Boris S. Mordukhovich, Dat Ba TranNeurIPS 2024 · 27 citations
