FlexIT: Towards Flexible Semantic Image Translation
Guillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk, Matthieu Cord
Abstract
Deep generative models, like GANs, have considerably improved the state of the art in image synthesis, and are able to generate near photo-realistic images in structured domains such as human faces. Based on this success, recent work on image editing proceeds by projecting images to the GAN latent space and manipulating the latent vector. However, these approaches are limited in that only images from a narrow domain can be transformed, and with only a limited number of editing operations. We propose FlexIT, a novel method which can take any input image and a user-defined text instruction for editing. Our method achieves flexible and natural editing, pushing the limits of semantic image translation. First, FlexIT combines the input image and text into a single target point in the CLIP multimodal embedding space. Via the latent space of an autoencoder, we iteratively transform the input image toward the target point, ensuring coherence and quality with a variety of novel regularization terms. We propose an evaluation protocol for semantic image translation, and thoroughly evaluate our method on ImageNet. Code will be available at https://github.com/facebookresearch/SemanticImageTranslation/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc05e05e-ac92-435a-a318-941c7c581c0fCited by top-tier papers13
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 102 citations
- PØDA: Prompt-driven Zero-shot Domain AdaptationMohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez et al.ICCV 2023 · 82 citations
- Diffusion-based Image Translation using disentangled style and content representationGihyun Kwon, Jong Chul YeICLR 2023 · 46 citations
- Towards Efficient Diffusion-Based Image Editing with Instant Attention MasksSiyu Zou, Jiji Tang, Yiyi Zhou, Jing He et al.AAAI 2024 · 24 citations
- FashionTex: Controllable Virtual Try-on with Text and TextureAnran Lin, Nanxuan Zhao, Shuliang Ning, Yuda Qiu et al.SIGGRAPH 2023 · 17 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
Related papers
- One Model to Edit Them All: Free-Form Text-Driven Image Manipulation with Semantic ModulationsYiming Zhu, Hongyu Liu, Yibing Song, Ziyang Yuan et al.NeurIPS 2022 · 44 citations
- Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image ManipulationXiwen Wei, Zhen Xu, Cheng Liu, Si Wu et al.CVPR 2023
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit DirectionsRameen Abdal, Peihao Zhu, John Femiani, Niloy J. Mitra et al.SIGGRAPH 2022 · 76 citations
- DeltaEdit: Exploring Text-free Training for Text-Driven Image ManipulationCVPR 2023
