Controllable and Compositional Generation with Latent-Space Energy-Based Models
Weili Nie, Arash Vahdat, Anima Anandkumar
摘要
Controllable generation is one of the key requirements for successful adoption of deep generative models in real-world applications, but it still remains as a great challenge. In particular, the compositional ability to generate novel concept combinations is out of reach for most current models. In this work, we use energybased models (EBMs) to handle compositional generation over a set of attributes. To make them scalable to high-resolution image generation, we introduce an EBM in the latent space of a pre-trained generative model such as StyleGAN. We propose a novel EBM formulation representing the joint distribution of data and attributes together, and we show how sampling from it is formulated as solving an ordinary differential equation (ODE). Given a pre-trained generator, all we need for controllable generation is to train an attribute classifier. Sampling with ODEs is done efficiently in the latent space and is robust to hyperparameters. Thus, our method is simple, fast to train, and efficient to sample. Experimental results show that our method outperforms the state-of-the-art in both conditional sampling and sequential editing. In compositional generation, our method excels at zero-shot generation of unseen attribute combinations. Also, by composing energy functions with logical operators, this work is the first to achieve such compositionality in generating photo-realistic images of resolution 1024×1024. Code is available at https://github.com/NVlabs/LACE . Recent works have attempted to overcome these issues by first training an unconditional generator, and then converting it to a conditional model with a small cost [26, 20, 1, 42] . This is often achieved by discovering semantically meaningful directions in the latent space of the unconditional model. This way, one would pay the most computational cost only once for training the unconditional model. However, these approaches often struggle with compositional generation, in particular with rare 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 被引用 903 次
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia 等ICLR 2024 · 被引用 161 次
- A Latent Space of Stochastic Diffusion Models for Zero-Shot Image Editing and GuidanceChen Henry Wu, Fernando De la TorreICCV 2023 · 被引用 141 次
- RoboDreamer: Learning Compositional World Models for Robot ImaginationSiyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li 等ICML 2024 · 被引用 140 次
- Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement LearningCheng Lu, Huayu Chen, Jianfei Chen, Hang Su 等ICML 2023 · 被引用 136 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
相关 Paper
- Everything is There in Latent Space: Attribute Editing and Attribute Style Manipulation by StyleGAN Latent Space ExplorationRishubh Parihar, Ankit Dhiman, Tejan Karmali, Venkatesh Babu R.ACM MM 2022 · 被引用 21 次
- VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsMoayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal 等ICCV 2023 · 被引用 3 次
- Latent Transformations via NeuralODEs for GAN-based Image EditingValentin Khrulkov, Leyla Mirvakhabova, Ivan V. Oseledets, Artem BabenkoICCV 2021 · 被引用 15 次
- Generative Visual Prompt: Unifying Distributional Control of Pre-Trained Generative ModelsChen Henry Wu, Saman Motamed, Shaunak Srivastava, Fernando De la TorreNeurIPS 2022 · 被引用 43 次
- Energy-Based Cross Attention for Bayesian Context Update in Text-to-Image Diffusion ModelsGeon Yeong Park, Jeongsol Kim, Beomsu Kim, Sang Wan Lee 等NeurIPS 2023 · 被引用 37 次
