Optimizing Information-theoretical Generalization Bound via Anisotropic Noise of SGLD
Bohan Wang, Huishuai Zhang, Jieyu Zhang, Qi Meng, Wei Chen, Tie-Yan Liu
Abstract
Recently, the information-theoretical framework has been proven to be able to obtain non-vacuous generalization bounds for large models trained by Stochastic Gradient Langevin Dynamics (SGLD) with isotropic noise. In this paper, we optimize the information-theoretical generalization bound by manipulating the noise structure in SGLD. We prove that with constraint to guarantee low empirical risk, the optimal noise covariance is the square root of the expected gradient covariance if both the prior and the posterior are jointly optimized. This validates that the optimal noise is quite close to the empirical gradient covariance. Technically, we develop a new information-theoretical bound that enables such an optimization analysis. We then apply matrix analysis to derive the form of optimal noise covariance. Presented constraint and results are validated by the empirical observations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 866aeba3-290d-48f1-afd0-a469245e447cCited by top-tier papers5
- Stability Based Generalization Bounds for Exponential Family Langevin DynamicsArindam Banerjee, Tiancong Chen, Xinyan Li, Yingxue ZhouICML 2022 · 9 citations
- Generalization Bounds via Conditional f-InformationZiqiao Wang, Yongyi MaoNeurIPS 2024 · 4 citations
- Beyond Lipschitz: Sharp Generalization and Excess Risk Bounds for Full-Batch GDKonstantinos E. Nikolakakis, Farzin Haddadpour, Amin Karbasi, Dionysios S. KalogeriasICLR 2023 · 3 citations
- Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGDZe Peng, Jian Zhang, Yisen Wang, Lei Qi et al.ICLR 2025
- Black-Box Generalization: Stability of Zeroth-Order LearningKonstantinos E. Nikolakakis, Farzin Haddadpour, Dionysios S. Kalogerias, Amin KarbasiNeurIPS 2022
Builds on5
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Sharpened Generalization Bounds based on Conditional Mutual Information and an Application to Noisy, Iterative AlgorithmsMahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M. Roy et al.NeurIPS 2020 · 124 citations
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 95 citations
- Conditioning and Processing: Techniques to Improve Information-Theoretic Generalization BoundsHassan Hafez-Kolahi, Zeinab Golgooni, Shohreh Kasaei, Mahdieh SoleymaniNeurIPS 2020 · 63 citations
- Sharper Generalization Bounds for Learning with Gradient-dominated Objective FunctionsYunwen Lei, Yiming YingICLR 2021 · 52 citations
Related papers
- Time-Independent Information-Theoretic Generalization Bounds for SGLDFutoshi Futami, Masahiro FujisawaNeurIPS 2023 · 12 citations
- Analyzing the Generalization Capability of SGLD Using Properties of Gaussian ChannelsHao Wang, Yizhe Huang, Rui Gao, Flávio P. CalmonNeurIPS 2021 · 32 citations
- Generalization of noisy SGD in unbounded non-convex settingsLeello Tadesse Dadi, Volkan CevherICML 2025
- Flatness-Aware Stochastic Gradient Langevin DynamicsStefano Bruno, Youngsik Hwang, JaeHyeon An, Sotirios Sabanis et al.ICML 2026
- Time-independent Generalization Bounds for SGLD in Non-convex SettingsTyler Farghly, Patrick RebeschiniNeurIPS 2021 · 30 citations
