Describe What to Change: A Text-guided Unsupervised Image-to-image Translation Approach
Yahui Liu, Marco De Nadai, Deng Cai, Huayang Li, Xavier Alameda-Pineda, Nicu Sebe, Bruno Lepri
Abstract
Manipulating visual attributes of images through human-written text is a very challenging task. On the one hand, models have to learn the manipulation without the ground truth of the desired output. On the other hand, models have to deal with the inherent ambiguity of natural language. Previous research usually requires either the user to describe all the characteristics of the desired image or to use richly-annotated image captioning datasets. In this work, we propose a novel unsupervised approach, based on image-to-image translation, that alters the attributes of a given image through a command-like sentence such as "change the hair color to black". Contrarily to state-of-the-art approaches, our model does not require a human-annotated dataset nor a textual description of all the attributes of the desired image, but only those that have to be modified. Our proposed model disentangles the image content from the visual attributes, and it learns to modify the latter using the textual description, before generating a new image from the content and the modified attribute representation. Because text might be inherently ambiguous (blond hair may refer to different shadows of blond, e.g. golden, icy, sandy), our method generates multiple stochastic versions of the same translation. Experiments show that the proposed model achieves promising performances on two large-scale public datasets: CelebA and CUB. We believe our approach will pave the way to new avenues of research combining textual and speech commands with visual attributes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy et al.ICCV 2021 · 162 citations
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 76 citations
- Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided DiffusionNisha Huang, Fan Tang, Weiming Dong, Changsheng XuACM MM 2022 · 49 citations
- Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language ModelZipeng Xu, Tianwei Lin, Hao Tang, Fu Li et al.CVPR 2022 · 38 citations
Builds on4
- SC-FEGAN: Face Editing Generative Adversarial Network With User's Sketch and ColorYoungjoo Jo, Jongyoul ParkICCV 2019 · 325 citations
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm et al.ICCV 2019 · 128 citations
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
- StarGAN v2: Diverse Image Synthesis for Multiple DomainsYunjey Choi, Youngjung Uh, Jaejun Yoo, Jung-Woo HaCVPR 2020
Related papers
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao et al.CVPR 2022 · 94 citations
- Retrieval Guided Unsupervised Multi-domain Image to Image TranslationRaul Gomez, Yahui Liu, Marco De Nadai, Dimosthenis Karatzas et al.ACM MM 2020 · 7 citations
- Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image ManipulationXiwen Wei, Zhen Xu, Cheng Liu, Si Wu et al.CVPR 2023
- HairCLIPv2: Unifying Hair Editing via Proxy Feature BlendingTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao et al.ICCV 2023 · 23 citations
- LANIT: Language-Driven Image-to-Image Translation for Unlabeled DataJihye Park, Sunwoo Kim, Soohyun Kim, Seokju Cho et al.CVPR 2023
