Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels
Hao Wang, Yizhe Huang, Rui Gao, Flávio P. Calmon
摘要
Optimization is a key component for training machine learning models and has a strong impact on their generalization. In this paper, we consider a particular optimization method -- the stochastic gradient Langevin dynamics (SGLD) algorithm -- and investigate the generalization of models trained by SGLD. We derive a new generalization bound by connecting SGLD with Gaussian channels found in information and communication theory. Our bound can be computed from the training data and incorporates the variance of gradients for quantifying a particular kind of sharpness of the loss landscape. We also consider a closely related algorithm with SGLD, namely differentially private SGD (DP-SGD). We prove that the generalization capability of DP-SGD can be amplified by iteration. Specifically, our bound can be sharpened by including a time-decaying factor if the DP-SGD algorithm outputs the last iterate while keeping other iterates hidden. This decay factor enables the contribution of early iterations to our bound to reduce with time and is established by strong data processing inequalities -- a fundamental tool in information theory. We demonstrate our bound through numerical experiments, showing that it can predict the behavior of the true generalization gap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and ImplicationsZiqiao Wang, Yongyi MaoICLR 2022 · 被引用 33 次
- A New Family of Generalization Bounds Using Samplewise Evaluated CMIFredrik Hellström, Giuseppe DurisiNeurIPS 2022 · 被引用 32 次
- Rate-Distortion Theoretic Bounds on Generalization Error for Distributed LearningMilad Sefidgaran, Romain Chor, Abdellatif ZaidiNeurIPS 2022 · 被引用 24 次
- Tighter Information-Theoretic Generalization Bounds from SupersamplesZiqiao Wang, Yongyi MaoICML 2023 · 被引用 23 次
- Time-Independent Information-Theoretic Generalization Bounds for SGLDFutoshi Futami, Masahiro FujisawaNeurIPS 2023 · 被引用 12 次
它引用的顶会 Paper4
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Sharpened Generalization Bounds based on Conditional Mutual Information and an Application to Noisy, Iterative AlgorithmsMahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M. Roy 等NeurIPS 2020 · 被引用 124 次
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 被引用 95 次
相关 Paper
- Generalization of noisy SGD in unbounded non-convex settingsLeello Tadesse Dadi, Volkan CevherICML 2025
- Privacy of Noisy Stochastic Gradient Descent: More Iterations without More Privacy LossJason M. Altschuler, Kunal TalwarNeurIPS 2022 · 被引用 89 次
- Optimizing Information-theoretical Generalization Bound via Anisotropic Noise of SGLDBohan Wang, Huishuai Zhang, Jieyu Zhang, Qi Meng 等NeurIPS 2021 · 被引用 3 次
- Stability Based Generalization Bounds for Exponential Family Langevin DynamicsArindam Banerjee, Tiancong Chen, Xinyan Li, Yingxue ZhouICML 2022 · 被引用 9 次
- Characterizing Membership Privacy in Stochastic Gradient Langevin DynamicsBingzhe Wu, Chaochao Chen, Shiwan Zhao, Cen Chen 等AAAI 2020 · 被引用 23 次
