Stochastic variance-reduced Gaussian variational inference on the Bures-Wasserstein manifold
Hoang Phuc Hau Luu, Hanlin Yu, Bernardo Williams, Marcelo Hartmann, Arto Klami
Abstract
Optimization in the Bures-Wasserstein space has been gaining popularity in the machine learning community since it draws connections between variational inference and Wasserstein gradient flows. The variational inference objective function of Kullback-Leibler divergence can be written as the sum of the negative entropy and the potential energy, making forward-backward Euler the method of choice. Notably, the backward step admits a closed-form solution in this case, facilitating the practicality of the scheme. However, the forward step is not exact since the Bures-Wasserstein gradient of the potential energy involves "intractable" expectations. Recent approaches propose using the Monte Carlo method -in practice a single-sample estimator -to approximate these terms, resulting in high variance and poor performance. We propose a novel variance-reduced estimator based on the principle of control variates. We theoretically show that this estimator has a smaller variance than the Monte-Carlo estimator in scenarios of interest. We also prove that variance reduction helps improve the optimization bounds of the current analysis. We empirically demonstrate that the proposed estimator gains order-of-magnitude improvements over previous Bures-Wasserstein methods. Recently, there has been emerging interest in Gaussian VI with a new geometric Riemannian optimization perspective (Lambert et al., 2022; Diao et al., 2023). The family of non-degenerate Gaus-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Large-Scale Wasserstein Gradient FlowsPetr Mokrov, Alexander Korotin, Lingxiao Li, Aude Genevay et al.NeurIPS 2021 · 112 citations
- The Wasserstein Proximal Gradient AlgorithmAdil Salim, Anna Korba, Giulia LuiseNeurIPS 2020 · 74 citations
- Averaging on the Bures-Wasserstein manifold: dimension-free convergence of gradient descentJason M. Altschuler, Sinho Chewi, Patrik Gerber, Austin J. StrommeNeurIPS 2021 · 60 citations
- Forward-Backward Gaussian Variational Inference via JKO in the Bures-Wasserstein SpaceMichael Ziyang Diao, Krishna Balasubramanian, Sinho Chewi, Adil SalimICML 2023 · 47 citations
- Provable convergence guarantees for black-box variational inferenceJustin Domke, Robert M. Gower, Guillaume GarrigosNeurIPS 2023 · 35 citations
Related papers
- Variational inference via Wasserstein gradient flowsMarc Lambert, Sinho Chewi, Francis R. Bach, Silvère Bonnabel et al.NeurIPS 2022 · 123 citations
- Sliced Wasserstein Estimation with Control VariatesKhai Nguyen, Nhat HoICLR 2024 · 16 citations
- On the Convergence of Projected Bures-Wasserstein Gradient Descent under Euclidean Strong ConvexityJunyi Fan, Yuxuan Han, Zijian Liu, Jian-Feng Cai et al.ICML 2024 · 2 citations
- Approximation Based Variance Reduction for Reparameterization GradientsTomas Geffner, Justin DomkeNeurIPS 2020 · 13 citations
- Variational inference via Gaussian interacting particles in the Bures-Wasserstein geometryGiacomo Borghi, Jose CarrilloICML 2026 · 2 citations
