Text as Neural Operator: Image Manipulation by Text Instruction
Tianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang, Honglak Lee, Irfan Essa
Abstract
n recent years, text-guided image manipulation has gained increasing attention in the multimedia and computer vision community. The input to conditional image generation has evolved from image-only to multimodality. In this paper, we study a setting that allows users to edit an image with multiple objects using complex text instructions to add, remove, or change the objects. The inputs of the task are multimodal including (1) a reference image and (2) an instruction in natural language that describes desired modifications to the image. We propose a GAN-based method to tackle this problem. The key idea is to treat text as neural operators to locally modify the image feature. We show that the proposed model performs favorably against recent strong baselines on three public datasets. Specifically, it generates images of greater fidelity and semantic relevance, and when used as a image query, leads to better retrieval performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b5b1e49-8809-4b02-aad3-947c9ac89cb5Cited by top-tier papers9
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy et al.ICCV 2021 · 162 citations
- TiGAN: Text-Based Interactive Image Generation and ManipulationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Chris Tensmeyer et al.AAAI 2022 · 18 citations
- ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-wise Semantic Alignment and GenerationJianan Wang, Guansong Lu, Hang Xu, Zhenguo Li et al.CVPR 2022 · 15 citations
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang et al.ACM MM 2022 · 12 citations
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 3 citations
Builds on14
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang et al.ICCV 2019 · 148 citations
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm et al.ICCV 2019 · 128 citations
- Diverse Image Synthesis From Semantic Layouts via Conditional IMLEKe Li, Tianhao Zhang, Jitendra MalikICCV 2019 · 102 citations
- Sequential Attention GAN for Interactive Image EditingYu Cheng, Zhe Gan, Yitong Li, Jingjing Liu et al.ACM MM 2020 · 74 citations
- Describe What to Change: A Text-guided Unsupervised Image-to-image Translation ApproachYahui Liu, Marco De Nadai, Deng Cai, Huayang Li et al.ACM MM 2020 · 58 citations
Related papers
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningZhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai et al.ACM MM 2020 · 17 citations
- UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text EditingLichen Ma, Xiaolong Fu, Gaojing Zhou, Zipeng Guo et al.AAAI 2026 · 1 citation
- CoSMo: Content-Style Modulation for Image Retrieval With Text FeedbackSeungmin Lee, Dongwan Kim, Bohyung HanCVPR 2021
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
- FlexIT: Towards Flexible Semantic Image TranslationGuillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk et al.CVPR 2022 · 36 citations
