Lightweight Generative Adversarial Networks for Text-Guided Image Manipulation
Bowen Li, Xiaojuan Qi, Philip H. S. Torr, Thomas Lukasiewicz
Abstract
We propose a novel lightweight generative adversarial network for efficient image manipulation using natural language descriptions. To achieve this, a new word-level discriminator is proposed, which provides the generator with fine-grained training feedback at word-level, to facilitate training a lightweight generator that has a small number of parameters, but can still correctly focus on specific visual attributes of an image, and then edit them without affecting other contents that are not described in the text. Furthermore, thanks to the explicit training signal related to each word, the discriminator can also be simplified to have a lightweight structure. Compared with the state of the art, our method has a much smaller number of parameters, but still achieves a competitive manipulation performance. Extensive experimental results demonstrate that our method can better disentangle different visual attributes, then correctly map them to corresponding semantic words, and thus achieve a more accurate image modification using natural language descriptions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Dual Contrastive Loss and Attention for GANsNing Yu, Guilin Liu, Aysegul Dundar, Andrew Tao et al.ICCV 2021 · 69 citations
- Language-based Photo Color Adjustment for Graphic DesignsZhenwei Wang, Nanxuan Zhao, Gerhard P. Hancke, Rynson W. H. LauSIGGRAPH 2023 · 49 citations
- Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language ModelZipeng Xu, Tianwei Lin, Hao Tang, Fu Li et al.CVPR 2022 · 38 citations
- DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image GenerationMengqi Huang, Zhendong Mao, Penghui Wang, Quan Wang et al.ACM MM 2022 · 26 citations
- DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksMing Tao, Bing-Kun Bao, Hao Tang, Fei Wu et al.AAAI 2023 · 19 citations
Builds on2
Related papers
- Uncovering the Disentanglement Capability in Text-to-Image Diffusion ModelsQiucheng Wu, Yujian Liu, Handong Zhao, Ajinkya Kale et al.CVPR 2023
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningZhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai et al.ACM MM 2020 · 17 citations
- Robust Conditional GAN from Uncertainty-Aware Pairwise ComparisonsLigong Han, Ruijiang Gao, Mun Kim, Xin Tao et al.AAAI 2020 · 14 citations
- Interpretable Generative Adversarial NetworksChao Li, Kelu Yao, Jin Wang, Boyu Diao et al.AAAI 2022 · 19 citations
- Everything is There in Latent Space: Attribute Editing and Attribute Style Manipulation by StyleGAN Latent Space ExplorationRishubh Parihar, Ankit Dhiman, Tejan Karmali, Venkatesh Babu R.ACM MM 2022 · 21 citations
