Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
Konstantinos Pitas
Abstract
Explaining how overparametrized neural networks simultaneously achieve low risk and zero empirical risk on benchmark datasets is an open problem. PAC-Bayes bounds optimized using variational inference (VI) have been recently proposed as a promising direction in obtaining non-vacuous bounds. We show empirically that this approach gives negligible gains when modeling the posterior as a Gaussian with diagonal covariance--known as the mean-field approximation. We investigate common explanations, such as the failure of VI due to problems in optimization or choosing a suboptimal prior. Our results suggest that investigating richer posteriors is the most promising direction forward.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1f789ab-849b-436d-8756-a7e8a60877e3Cited by top-tier papers2
- On the generalization of learning algorithms that do not convergeNisha Chandramoorthy, Andreas Loukas, Khashayar Gatmiry, Stefanie JegelkaNeurIPS 2022 · 13 citations
- Optimization and Bayes: A Trade-off for Overparameterized Neural NetworksZhengmian Hu, Heng HuangNeurIPS 2023 · 1 citation
Builds on4
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian AnalysisYusuke Tsuzuku, Issei Sato, Masashi SugiyamaICML 2020 · 91 citations
- In Defense of Uniform Convergence: Generalization via Derandomization with an Application to Interpolating PredictorsJeffrey Negrea, Gintare Karolina Dziugaite, Daniel M. RoyICML 2020 · 66 citations
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
- Information-Theoretic Local Minima Characterization and RegularizationZhiwei Jia, Hao SuICML 2020 · 22 citations
Related papers
- Efficient Low Rank Gaussian Variational Inference for Neural NetworksMarcin Tomczak, Siddharth Swaroop, Richard E. TurnerNeurIPS 2020 · 37 citations
- On the detrimental effect of invariances in the likelihood for variational inferenceRichard Kurle, Ralf Herbrich, Tim Januschowski, Yuyang Wang et al.NeurIPS 2022 · 10 citations
- VIKING: Deep variational inference with stochastic projectionsSamuel Matthiesen, Hrittik Roy, Nicholas Krämer, Yevgen Zainchkovskyy et al.NeurIPS 2025 · 3 citations
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural NetworksJakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling, Linh Tran et al.ICML 2020 · 52 citations
- Collapsed Variational Bounds for Bayesian Neural NetworksMarcin Tomczak, Siddharth Swaroop, Andrew Y. K. Foong, Richard E. TurnerNeurIPS 2021 · 14 citations
