Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels
Hao Wang, Yizhe Huang, Rui Gao, Flávio P. Calmon
Abstract
Optimization is a key component for training machine learning models and has a strong impact on their generalization. In this paper, we consider a particular optimization method -- the stochastic gradient Langevin dynamics (SGLD) algorithm -- and investigate the generalization of models trained by SGLD. We derive a new generalization bound by connecting SGLD with Gaussian channels found in information and communication theory. Our bound can be computed from the training data and incorporates the variance of gradients for quantifying a particular kind of sharpness of the loss landscape. We also consider a closely related algorithm with SGLD, namely differentially private SGD (DP-SGD). We prove that the generalization capability of DP-SGD can be amplified by iteration. Specifically, our bound can be sharpened by including a time-decaying factor if the DP-SGD algorithm outputs the last iterate while keeping other iterates hidden. This decay factor enables the contribution of early iterations to our bound to reduce with time and is established by strong data processing inequalities -- a fundamental tool in information theory. We demonstrate our bound through numerical experiments, showing that it can predict the behavior of the true generalization gap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and ImplicationsZiqiao Wang, Yongyi MaoICLR 2022 · 33 citations
- A New Family of Generalization Bounds Using Samplewise Evaluated CMIFredrik Hellström, Giuseppe DurisiNeurIPS 2022 · 32 citations
- Rate-Distortion Theoretic Bounds on Generalization Error for Distributed LearningMilad Sefidgaran, Romain Chor, Abdellatif ZaidiNeurIPS 2022 · 24 citations
- Tighter Information-Theoretic Generalization Bounds from SupersamplesZiqiao Wang, Yongyi MaoICML 2023 · 23 citations
- Time-Independent Information-Theoretic Generalization Bounds for SGLDFutoshi Futami, Masahiro FujisawaNeurIPS 2023 · 12 citations
Builds on4
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Sharpened Generalization Bounds based on Conditional Mutual Information and an Application to Noisy, Iterative AlgorithmsMahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M. Roy et al.NeurIPS 2020 · 124 citations
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 95 citations
Related papers
- Generalization of noisy SGD in unbounded non-convex settingsLeello Tadesse Dadi, Volkan CevherICML 2025
- Privacy of Noisy Stochastic Gradient Descent: More Iterations without More Privacy LossJason M. Altschuler, Kunal TalwarNeurIPS 2022 · 89 citations
- Optimizing Information-theoretical Generalization Bound via Anisotropic Noise of SGLDBohan Wang, Huishuai Zhang, Jieyu Zhang, Qi Meng et al.NeurIPS 2021 · 3 citations
- Stability Based Generalization Bounds for Exponential Family Langevin DynamicsArindam Banerjee, Tiancong Chen, Xinyan Li, Yingxue ZhouICML 2022 · 9 citations
- Characterizing Membership Privacy in Stochastic Gradient Langevin DynamicsBingzhe Wu, Chaochao Chen, Shiwan Zhao, Cen Chen et al.AAAI 2020 · 23 citations
