On the Value of Infinite Gradients in Variational Autoencoder Models
Bin Dai, Wenliang Li, David P. Wipf
Abstract
A number of recent studies of continuous variational autoencoder (VAE) models have noted, either directly or indirectly, the tendency of various parameter gradients to drift towards infinity during training. Because such gradients could potentially contribute to numerical instabilities, and are often framed as a problematic phenomena to be avoided, it may be tempting to shift to alternative energy functions that guarantee bounded gradients. But it remains an open question: What might the unintended consequences of such a restriction be? To address this issue, we examine how unbounded gradients relate to the regularization of a broad class of autoencoder-based architectures, including VAE models, as applied to data lying on or near a low-dimensional manifold (e.g., natural images). Our main finding is that, if the ultimate goal is to simultaneously avoid over-regularization (high reconstruction errors, sometimes referred to as posterior collapse) and underregularization (excessive latent dimensions are not pruned from the model), then an autoencoder-based energy function with infinite gradients around optimal representations is provably required per a certain technical sense which we carefully detail. Given that both over-and under-regularization can directly lead to poor generated sample quality or suboptimal feature selection, this result suggests that heuristic modifications to or constraints on the VAE energy function may at times be ill-advised, and large gradients should be accommodated to the extent possible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b448a88-4893-476b-890f-422a5cac979cCited by top-tier papers3
- Learning Manifold Dimensions with Conditional Variational AutoencodersYijia Zheng, Tong He, Yixuan Qiu, David P. WipfNeurIPS 2022 · 34 citations
- Marginalization is not Marginal: No Bad VAE Local Minima when Learning Optimal Sparse RepresentationsDavid WipfICML 2023 · 5 citations
- Sparse Autoencoders, Again?Yin Lu, Xuening Zhu, Tong He, David WipfICML 2025
Builds on3
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- From Variational to Deterministic AutoencodersPartha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael J. Black et al.ICLR 2020 · 298 citations
- The Usual Suspects? Reassessing Blame for VAE Posterior CollapseBin Dai, Ziyu Wang, David P. WipfICML 2020 · 89 citations
Related papers
- Autoencoder Image Interpolation by Shaping the Latent SpaceAlon Oring, Zohar Yakhini, Yacov Hel-OrICML 2021 · 41 citations
- Iterative energy-based projection on a normal data manifold for anomaly localizationDavid Dehaene, Oriel Frigo, Sébastien Combrexelle, Pierre ElineICLR 2020 · 157 citations
- VAE Approximation Error: ELBO and Exponential FamiliesAlexander Shekhovtsov, Dmitrij Schlesinger, Boris FlachICLR 2022 · 21 citations
- Improving Variational Autoencoders with Density Gap-based RegularizationJianfei Zhang, Jun Bai, Chenghua Lin, Yanmeng Wang et al.NeurIPS 2022 · 11 citations
- Variational autoencoders in the presence of low-dimensional data: landscape and implicit biasFrederic Koehler, Viraj Mehta, Chenghui Zhou, Andrej RisteskiICLR 2022 · 14 citations
