ManiGAN: Text-Guided Image Manipulation
Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. Torr
Abstract
The goal of our paper is to semantically edit parts of an image to match a given text that describes desired attributes (e.g., texture, colour, and background), while preserving other contents that are irrelevant to the text. To achieve this, we propose a novel generative adversarial network (ManiGAN), which contains two key components: text-image affine combination module (ACM) and detail correction module (DCM). The ACM selects image regions relevant to the given text and then correlates the regions with corresponding semantic words for effective manipulation. Meanwhile, it encodes original image features to help reconstruct text-irrelevant contents. The DCM rectifies mismatched attributes and completes missing contents of the synthetic image. Finally, we suggest a new metric for evaluating image manipulation results, in terms of both the generation of new attributes and the reconstruction of text-irrelevant contents. Extensive experiments on the CUB and COCO datasets demonstrate the superior performance of the proposed method. Code is available at https://github.com/mrlibw/ManiGAN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers72
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan et al.ICCV 2023 · 770 citations
- CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsCan Wang, Menglei Chai, Mingming He, Dongdong Chen et al.CVPR 2022 · 313 citations
- CLIPstyler: Image Style Transfer with a Single Text ConditionGihyun Kwon, Jong Chul YeCVPR 2022 · 224 citations
- ObjectFormer for Image Manipulation Detection and LocalizationJunke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han et al.CVPR 2022 · 190 citations
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
Builds on2
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm et al.ICCV 2019 · 128 citations
- Local Class-Specific and Global Image-Level Generative Adversarial Networks for Semantic-Guided Scene GenerationHao Tang, Dan Xu, Yan Yan, Philip H. S. Torr et al.CVPR 2020
Related papers
- ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-wise Semantic Alignment and GenerationJianan Wang, Guansong Lu, Hang Xu, Zhenguo Li et al.CVPR 2022 · 15 citations
- Lightweight Generative Adversarial Networks for Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Philip H. S. Torr, Thomas LukasiewiczNeurIPS 2020 · 76 citations
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 3 citations
- SAT3D: Image-driven Semantic Attribute Transfer in 3DZhijun Zhai, Zengmao Wang, Xiaoxiao Long, Kaixuan Zhou et al.ACM MM 2024
- Enjoy Your Editing: Controllable GANs for Image Editing via Latent Space NavigationPeiye Zhuang, Oluwasanmi Koyejo, Alexander G. SchwingICLR 2021 · 88 citations
