Understanding Why Neural Networks Generalize Well Through GSNR of Parameters
Jinlong Liu, Yunzhi Bai, Guoqing Jiang, Ting Chen, Huayan Wang
Abstract
As deep neural networks (DNNs) achieve tremendous success across many application domains, researchers tried to explore in many aspects on why they generalize well. In this paper, we provide a novel perspective on these issues using the gradient signal to noise ratio (GSNR) of parameters during training process of DNNs. The GSNR of a parameter is simply defined as the ratio between its gradient's squared mean and variance, over the data distribution. Based on several approximations, we establish a quantitative relationship between model parameters' GSNR and the generalization gap. This relationship indicates that larger GSNR during training process leads to better generalization performance. Futher, we show that, different from that of shallow models (e.g. logistic regression, support vector machines), the gradient descent optimization dynamics of DNNs naturally produces large GSNR during training, which is probably the key to DNNs’ remarkable generalization ability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- Dataset Condensation with Contrastive SignalsSaehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun et al.ICML 2022 · 139 citations
- KungFu: Making Training in Distributed Machine Learning AdaptiveLuo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis et al.OSDI 2020 · 92 citations
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou et al.NeurIPS 2021 · 51 citations
- Robust Optimization for Multilingual Translation with Imbalanced DataXian Li, Hongyu GongNeurIPS 2021 · 24 citations
Related papers
- How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?Wei Huang, Andi Han, Yujin Song, Yilan Chen et al.NeurIPS 2025 · 4 citations
- Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersMateusz Michalkiewicz, Masoud Faraki, Xiang Yu, Manmohan Chandraker et al.ICCV 2023 · 9 citations
- Sharp Generalization for Nonparametric Regression by Over-Parameterized Neural Networks: A Distribution-Free Analysis in Spherical CovariateYingzhen YangICML 2025
- Benign Oscillation of Stochastic Gradient Descent with Large Learning RateMiao Lu, Beining Wu, Xiaodong Yang, Difan ZouICLR 2024 · 9 citations
- Learning Trajectories are Generalization IndicatorsJingwen Fu, Zhizheng Zhang, Dacheng Yin, Yan Lu et al.NeurIPS 2023 · 6 citations
