StEP: Style-Based Encoder Pre-Training for Multi-Modal Image Synthesis
Moustafa Meshry, Yixuan Ren, Larry S. Davis, Abhinav Shrivastava
摘要
We propose a novel approach for multi-modal Imageto-image (I2I) translation. To tackle the one-to-many relationship between input and output domains, previous works use complex training objectives to learn a latent embedding, jointly with the generator, that models the variability of the output domain. In contrast, we directly model the style variability of images, independent of the image synthesis task. Specifically, we pre-train a generic style encoder using a novel proxy task to learn an embedding of images, from arbitrary domains, into a low-dimensional style latent space. The learned latent space introduces several advantages over previous traditional approaches to multi-modal I2I translation. First, it is not dependent on the target dataset, and generalizes well across multiple domains. Second, it learns a more powerful and expressive latent space, which improves the fidelity of style capture and transfer. The proposed style pre-training also simplifies the training objective and speeds up the training significantly. Furthermore, we provide a detailed study of the contribution of different loss terms to the task of multi-modal I2I translation, and propose a simple alternative to VAEs to enable sampling from unconstrained latent spaces. Finally, we achieve state-of-the-art results on six challenging benchmarks with a simple training objective that includes only a GAN loss and a reconstruction loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unsupervised Image-to-Image Translation with Generative PriorShuai Yang, Liming Jiang, Ziwei Liu, Chen Change LoyCVPR 2022 · 被引用 51 次
- Chop & Learn: Recognizing and Generating Object-State CompositionsNirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara 等ICCV 2023 · 被引用 20 次
它引用的顶会 Paper2
相关 Paper
- Encoding in Style: A StyleGAN Encoder for Image-to-Image TranslationElad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan 等CVPR 2021
- Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image TranslationYahui Liu, Enver Sangineto, Yajing Chen, Linchao Bao 等CVPR 2021
- Harnessing the Conditioning Sensorium for Improved Image TranslationCooper Nederhood, Nicholas I. Kolkin, Deqing Fu, Jason SalavonICCV 2021 · 被引用 6 次
- Unpaired Image-to-Image Translation via Latent Energy TransportYang Zhao, Changyou ChenCVPR 2021
- A Style-aware Discriminator for Controllable Image TranslationKunhee Kim, Sanghun Park, Eunyeong Jeon, Taehun Kim 等CVPR 2022 · 被引用 31 次
