Conditional Image Generation by Conditioning Variational Auto-Encoders
William Harvey, Saeid Naderiparizi, Frank Wood
Abstract
We present a conditional variational auto-encoder (VAE) which, to avoid the substantial cost of training from scratch, uses an architecture and training objective capable of leveraging a foundation model in the form of a pretrained unconditional VAE. To train the conditional VAE, we only need to train an artifact to perform amortized inference over the unconditional VAE's latent variables given a conditioning input. We demonstrate our approach on tasks including image inpainting, for which it outperforms state-of-the-art GAN-based approaches at faithfully representing the inherent uncertainty. We conclude by describing a possible application of our inpainting model, in which it is used to perform Bayesian experimental design for the purpose of guiding a sensor. * Frank Wood is also affiliated with the Montréal Institute for Learning Algorithms (Mila) and Inverted AI. Figure 1: Left column: Images with most pixels masked out. Rest: Completions from our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Dense Policy: Bidirectional Autoregressive Learning of ActionsYue Su, Xinyu Zhan, Hongjie Fang, Han Xue et al.ICCV 2025 · 23 citations
- Cross City Traffic Flow Generation via Retrieval Augmented Diffusion ModelYudong Li, Jingyuan Wang, Xie Yu, Peiyu Wang et al.NeurIPS 2025 · 9 citations
- VisDiff: SDF-Guided Polygon Generation for Visibility Reconstruction, Characterization and RecognitionRahul Moorthy Mahesh, Jun-Jee Chao, Volkan IslerNeurIPS 2025 · 1 citation
- One-shot Conditional Sampling: MMD meets Nearest NeighborsAnirban Chatterjee, Sayantan Choudhury, Rohan HoreICML 2026
- RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned PriorChing Hua Lee, Chouchang Yang, Jaejin Cho, Yashas Malur Saidutta et al.ICML 2025
Builds on15
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Large Scale Image Completion via Co-Modulated Generative Adversarial NetworksShengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong et al.ICLR 2021 · 348 citations
- High-Fidelity Pluralistic Image Completion with TransformersZiyu Wan, Jingbo Zhang, Dongdong Chen, Jing LiaoICCV 2021 · 296 citations
Related papers
- Meta-Learning with Shared Amortized Variational InferenceEkaterina Iakovleva, Jakob Verbeek, Karteek AlahariICML 2020 · 25 citations
- Generalization Gap in Amortized InferenceMingtian Zhang, Peter Hayes, David BarberNeurIPS 2022 · 14 citations
- The Autoencoding Variational AutoencoderA. Taylan Cemgil, Sumedh Ghaisas, Krishnamurthy Dvijotham, Sven Gowal et al.NeurIPS 2020 · 81 citations
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
- Variational Amodal Object CompletionHuan Ling, David Acuna, Karsten Kreis, Seung Wook Kim et al.NeurIPS 2020 · 56 citations
