NVAE: A Deep Hierarchical Variational Autoencoder
Arash Vahdat, Jan Kautz
Abstract
Normalizing flows, autoregressive models, variational autoencoders (VAEs), and deep energy-based models are among competing likelihood-based frameworks for deep generative learning. Among them, VAEs have the advantage of fast and tractable sampling and easy-to-access encoding networks. However, they are currently outperformed by other models such as normalizing flows and autoregressive models. While the majority of the research in VAEs is focused on the statistical challenges, we explore the orthogonal direction of carefully designing neural architectures for hierarchical VAEs. We propose Nouveau VAE (NVAE), a deep hierarchical VAE built for image generation using depth-wise separable convolutions and batch normalization. NVAE is equipped with a residual parameterization of Normal distributions and its training is stabilized by spectral regularization. We show that NVAE achieves state-of-the-art results among non-autoregressive likelihood-based models on the MNIST, CIFAR-10, CelebA 64, and CelebA HQ datasets and it provides a strong baseline on FFHQ. For example, on CIFAR-10, NVAE pushes the state-of-the-art from 2.98 to 2.91 bits per dimension, and it produces high-quality images on CelebA HQ as shown in Fig. 1 . To the best of our knowledge, NVAE is the first successful VAE applied to natural images as large as 256×256 pixels. The source code is available at https://github.com/NVlabs/NVAE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2267f205-9024-4191-a873-7aa494b271b8Cited by top-tier papers292
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Diffusion Probabilistic FieldsPeiye Zhuang, Samira Abnar, Jiatao Gu, Alexander G. Schwing et al.ICLR 2023 · 3,587 citations
Builds on2
Related papers
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on ImagesRewon ChildICLR 2021 · 45 citations
- Inverse problem regularization with hierarchical variational autoencodersJean Prost, Antoine Houdard, Andrés Almansa, Nicolas PapadakisICCV 2023 · 10 citations
- Hierarchical Quantized AutoencodersWill Williams, Sam Ringer, Tom Ash, David MacLeod et al.NeurIPS 2020 · 90 citations
- Bit Prioritization in Variational Autoencoders via Progressive CodingRui Shu, Stefano ErmonICML 2022 · 9 citations
- A Contrastive Learning Approach for Training Variational Autoencoder PriorsJyoti Aneja, Alexander G. Schwing, Jan Kautz, Arash VahdatNeurIPS 2021 · 112 citations
