Handling the Positive-Definite Constraint in the Bayesian Learning Rule
Wu Lin, Mark Schmidt, Mohammad Emtiyaz Khan
Abstract
The Bayesian learning rule is a natural-gradient variational inference method, which not only contains many existing learning algorithms as special cases but also enables the design of new algorithms. Unfortunately, when variational parameters lie in an open constraint set, the rule may not satisfy the constraint and requires line-searches which could slow down the algorithm. In this work, we address this issue for positive-definite constraints by proposing an improved rule that naturally handles the constraints. Our modification is obtained by using Riemannian gradient methods, and is valid when the approximation attains a block-coordinate natural parameterization (e.g., Gaussian distributions and their mixtures). We propose a principled way to derive Riemannian gradients and retractions from scratch. Our method outperforms existing methods without any significant increase in computation. Our work makes it easier to apply the rule in the presence of positive-definite constraints in parameter spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bc688ed-ad7b-42b2-ba06-b98d31c19e08Cited by top-tier papers13
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- Forward-Backward Gaussian Variational Inference via JKO in the Bures-Wasserstein SpaceMichael Ziyang Diao, Krishna Balasubramanian, Sinho Chewi, Adil SalimICML 2023 · 47 citations
- Tractable structured natural-gradient descent using local parameterizationsWu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark SchmidtICML 2021 · 36 citations
- Beyond Deep Ensembles: A Large-Scale Evaluation of Bayesian Deep Learning under Distribution ShiftFlorian Seligmann, Philipp Becker, Michael Volpp, Gerhard NeumannNeurIPS 2023 · 32 citations
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae et al.ICML 2024 · 23 citations
Related papers
- The Quotient Bayesian Learning RuleMykola Lukashchuk, Raphaël Trésor, Wouter W. L. Nuijten, Ismail Senöz et al.NeurIPS 2025 · 1 citation
- Riemannian Laplace approximations for Bayesian neural networksFederico Bergamin, Pablo Moreno-Muñoz, Søren Hauberg, Georgios ArvanitidisNeurIPS 2023 · 18 citations
- Adaptive gradient descent on Riemannian manifolds and its applications to Gaussian variational inferenceJiyoung Park, Jaewook J. Suh, Bofan Wang, Anirban Bhattacharya et al.ICLR 2026
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 10 citations
- Riemannian stochastic optimization methods avoid strict saddle pointsYa-Ping Hsieh, Mohammad Reza Karimi Jaghargh, Andreas Krause, Panayotis MertikopoulosNeurIPS 2023 · 17 citations
