On the Importance of Gradient Norm in PAC-Bayesian Bounds
Itai Gat, Yossi Adi, Alexander G. Schwing, Tamir Hazan
Abstract
Generalization bounds which assess the difference between the true risk and the empirical risk, have been studied extensively. However, to obtain bounds, current techniques use strict assumptions such as a uniformly bounded or a Lipschitz loss function. To avoid these assumptions, in this paper, we follow an alternative approach: we relax uniform bounds assumptions by using on-average bounded loss and on-average bounded gradient norm assumptions. Following this relaxation, we propose a new generalization bound that exploits the contractivity of the log-Sobolev inequalities. These inequalities add an additional loss-gradient norm term to the generalization bound, which is intuitively a surrogate of the model complexity. We apply the proposed bound on Bayesian deep nets and empirically analyze the effect of this new loss-gradient norm term on different neural architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f05ce2f6-bbe7-4050-83c6-a730c59cfa19Cited by top-tier papers3
- Layer Collaboration in the Forward-Forward AlgorithmGuy Lorberbom, Itai Gat, Yossi Adi, Alexander G. Schwing et al.AAAI 2024 · 22 citations
- CR-SAM: Curvature Regularized Sharpness-Aware MinimizationTao Wu, Tie Luo, Donald C. Wunsch IIAAAI 2024 · 15 citations
- PAC-Bayes-Chernoff bounds for unbounded lossesIoar Casado, Luis A. Ortega Andrés, Aritz Pérez, Andrés R. MasegosaNeurIPS 2024 · 14 citations
Builds on5
- Sharpened Generalization Bounds based on Conditional Mutual Information and an Application to Noisy, Iterative AlgorithmsMahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M. Roy et al.NeurIPS 2020 · 124 citations
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 95 citations
- Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform StabilityAlec Farid, Anirudha MajumdarNeurIPS 2021 · 46 citations
- Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-LearningNan Ding, Xi Chen, Tomer Levinboim, Sebastian Goodman et al.NeurIPS 2021 · 34 citations
- How Tight Can PAC-Bayes be in the Small Data Regime?Andrew Y. K. Foong, Wessel P. Bruinsma, David R. Burt, Richard E. TurnerNeurIPS 2021 · 28 citations
Related papers
- PAC-Bayesian Spectrally-Normalized Bounds for Adversarially Robust GeneralizationJiancong Xiao, Ruoyu Sun, Zhi-Quan LuoNeurIPS 2023 · 14 citations
- Estimating Lipschitz constants of monotone deep equilibrium modelsChirag Pabbaraju, Ezra Winston, J. Zico KolterICLR 2021 · 33 citations
- Approximation with CNNs in Sobolev Space: with Applications to ClassificationGuohao Shen, Yuling Jiao, Yuanyuan Lin, Jian HuangNeurIPS 2022 · 25 citations
- What training reveals about neural network complexityAndreas Loukas, Marinos Poiitis, Stefanie JegelkaNeurIPS 2021 · 12 citations
- Improved Sample Complexities for Deep Neural Networks and Robust Classification via an All-Layer MarginColin Wei, Tengyu MaICLR 2020 · 91 citations
