Unsupervised K-modal styled content generation
Omry Sendik, Dani Lischinski, Daniel Cohen-Or
Abstract
The emergence of deep generative models has recently enabled the automatic generation of massive amounts of graphical content, both in 2D and in 3D. Generative Adversarial Networks (GANs) and style control mechanisms, such as Adaptive Instance Normalization (AdaIN), have proved particularly effective in this context, culminating in the state-of-the-art StyleGAN architecture. While such models are able to learn diverse distributions, provided a sufficiently large training set, they are not well-suited for scenarios where the distribution of the training data exhibits a multi-modal behavior. In such cases, reshaping a uniform or normal distribution over the latent space into a complex multi-modal distribution in the data domain is challenging, and the generator might fail to sample the target distribution well. Furthermore, existing unsupervised generative models are not able to control the mode of the generated samples independently of the other visual attributes, despite the fact that they are typically disentangled in the training data. In this paper, we introduce uMM-GAN, a novel architecture designed to better model multi-modal distributions, in an unsupervised fashion. Building upon the StyleGAN architecture, our network learns multiple modes, in a completely unsupervised manner , and combines them using a set of learned weights. We demonstrate that this approach is capable of effectively approximating a complex distribution as a superposition of multiple simple ones. We further show that uMM-GAN effectively disentangles between modes and style, thereby providing an independent degree of control over the generated content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6e43d1f-2c74-4ba4-89da-4de8fea68a00Cited by top-tier papers2
- Designing an encoder for StyleGAN image manipulationOmer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik et al.SIGGRAPH 2021 · 692 citations
- Self-Distilled StyleGAN: Towards Generation from Internet PhotosRon Mokady, Omer Tov, Michal Yarom, Oran Lang et al.SIGGRAPH 2022 · 29 citations
Builds on1
Related papers
- Adversarial Latent AutoencodersStanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco DorettoCVPR 2020
- Multi-Class Multi-Instance Count Conditioned Adversarial Image GenerationAmrutha Saseendran, Kathrin Skubch, Margret KeuperICCV 2021 · 2 citations
- Image Synthesis From Reconfigurable Layout and StyleWei Sun, Tianfu WuICCV 2019 · 160 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- Analyzing and Improving the Image Quality of StyleGANTero Karras, Samuli Laine, Miika Aittala, Janne Hellsten et al.CVPR 2020
