On the detrimental effect of invariances in the likelihood for variational inference
Richard Kurle, Ralf Herbrich, Tim Januschowski, Yuyang Wang, Jan Gasthaus
Abstract
Variational Bayesian posterior inference often requires simplifying approximations such as mean-field parametrisation to ensure tractability. However, prior work has associated the variational mean-field approximation for Bayesian neural networks with underfitting in the case of small datasets or large model sizes. In this work, we show that invariances in the likelihood function of over-parametrised models contribute to this phenomenon because these invariances complicate the structure of the posterior by introducing discrete and/or continuous modes which cannot be well approximated by Gaussian mean-field distributions. In particular, we show that the mean-field approximation has an additional gap in the evidence lower bound compared to a purpose-built posterior that takes into account the known invariances. Importantly, this invariance gap is not constant; it vanishes as the approximation reverts to the prior. We proceed by first considering translation invariances in a linear model with a single data point in detail. We show that, while the true posterior can be constructed from a mean-field parametrisation, this is achieved only if the objective function takes into account the invariance gap. Then, we transfer our analysis of the linear model to neural networks. Our analysis provides a framework for future work to explore solutions to the invariance problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb595ae6-21a9-4ac2-a02f-6b620f8d09d9Cited by top-tier papers2
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- On permutation symmetries in Bayesian neural network posteriors: a variational perspectiveSimone Rossi, Ankit Singh, Thomas HannaganNeurIPS 2023 · 5 citations
Builds on11
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsMichael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma et al.ICML 2020 · 239 citations
- On the Expressiveness of Approximate Inference in Bayesian Neural NetworksAndrew Y. K. Foong, David R. Burt, Yingzhen Li, Richard E. TurnerNeurIPS 2020 · 142 citations
Related papers
- VIKING: Deep variational inference with stochastic projectionsSamuel Matthiesen, Hrittik Roy, Nicholas Krämer, Yevgen Zainchkovskyy et al.NeurIPS 2025 · 3 citations
- Reparameterization invariance in approximate Bayesian inferenceHrittik Roy, Marco Miani, Carl Henrik Ek, Philipp Hennig et al.NeurIPS 2024 · 20 citations
- Liberty or Depth: Deep Bayesian Neural Nets Do Not Need Complex Weight Posterior ApproximationsSebastian Farquhar, Lewis Smith, Yarin GalNeurIPS 2020 · 47 citations
- Collapsed Variational Bounds for Bayesian Neural NetworksMarcin Tomczak, Siddharth Swaroop, Andrew Y. K. Foong, Richard E. TurnerNeurIPS 2021 · 14 citations
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural NetworksJakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling, Linh Tran et al.ICML 2020 · 52 citations
