Marginalization is not Marginal: No Bad VAE Local Minima when Learning Optimal Sparse Representations
David Wipf
Abstract
Although the variational autoencoder (VAE) represents a widely-used deep generative model, the underlying energy function when applied to continuous data remains poorly understood. In fact, most prior theoretical analysis has assumed a simplified affine decoder such that the model collapses to probabilistic PCA, a restricted regime whereby existing classical algorithms can also be trivially applied to guarantee globally optimal solutions. To push our understanding into more complex, practically-relevant settings, this paper instead adopts a deceptively sophisticated single-layer decoder that nonetheless allows the VAE to address the fundamental challenge of learning optimally sparse representations of continuous data originating from popular multiple-response regression models. In doing so, we can then examine VAE properties within the non-trivial context of solving difficult, NP-hard inverse problems. More specifically, we prove rigorous conditions which guarantee that any minimum of the VAE energy (local or global) will produce the optimally sparse latent representation, meaning zero reconstruction error using a minimal number of active latent dimensions. This is ultimately possible because VAE marginalization over the latent posterior selectively smooths away bad local minima as has been conjectured but not actually proven in prior work. We then discuss how equivalent-capacity deterministic autoencoders, even with appropriate sparsity-promoting regularization of the latent space, maintain bad local minima that do not correspond with such parsimonious representations. Overall, these results serve to elucidate key properties of the VAE loss surface relative to finding low-dimensional structure in data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f886ff92-c11b-442f-9a2e-572c240bd4ffCited by top-tier papers2
- Learning Manifold Dimensions with Conditional Variational AutoencodersYijia Zheng, Tong He, Yixuan Qiu, David P. WipfNeurIPS 2022 · 34 citations
- Be a Goldfish: Forgetting Bad Conditioning in Sparse Linear Regression via Variational AutoencodersKuheli Pratihar, Debdeep MukhopadhyayICML 2025
Builds on5
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Learning Manifold Dimensions with Conditional Variational AutoencodersYijia Zheng, Tong He, Yixuan Qiu, David P. WipfNeurIPS 2022 · 34 citations
- VAE Approximation Error: ELBO and Exponential FamiliesAlexander Shekhovtsov, Dmitrij Schlesinger, Boris FlachICLR 2022 · 21 citations
- On the Value of Infinite Gradients in Variational Autoencoder ModelsBin Dai, Wenliang Li, David P. WipfNeurIPS 2021 · 15 citations
- Variational autoencoders in the presence of low-dimensional data: landscape and implicit biasFrederic Koehler, Viraj Mehta, Chenghui Zhou, Andrej RisteskiICLR 2022 · 14 citations
Related papers
- The Usual Suspects? Reassessing Blame for VAE Posterior CollapseBin Dai, Ziyu Wang, David P. WipfICML 2020 · 89 citations
- Sparse Autoencoders, Again?Yin Lu, Xuening Zhu, Tong He, David WipfICML 2025
- Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic AutoencodersAmrutha Saseendran, Kathrin Skubch, Stefan Falkner, Margret KeuperNeurIPS 2021 · 13 citations
- Controlling Posterior Collapse by an Inverse Lipschitz Constraint on the Decoder NetworkYuri Kinoshita, Kenta Oono, Kenji Fukumizu, Yuichi Yoshida et al.ICML 2023 · 6 citations
- Posterior Collapse and Latent Variable Non-identifiabilityYixin Wang, David M. Blei, John P. CunninghamNeurIPS 2021 · 97 citations
