Text as Neural Operator: Image Manipulation by Text Instruction
Tianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang, Honglak Lee, Irfan Essa
摘要
n recent years, text-guided image manipulation has gained increasing attention in the multimedia and computer vision community. The input to conditional image generation has evolved from image-only to multimodality. In this paper, we study a setting that allows users to edit an image with multiple objects using complex text instructions to add, remove, or change the objects. The inputs of the task are multimodal including (1) a reference image and (2) an instruction in natural language that describes desired modifications to the image. We propose a GAN-based method to tackle this problem. The key idea is to treat text as neural operators to locally modify the image feature. We show that the proposed model performs favorably against recent strong baselines on three public datasets. Specifically, it generates images of greater fidelity and semantic relevance, and when used as a image query, leads to better retrieval performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy 等ICCV 2021 · 被引用 162 次
- TiGAN: Text-Based Interactive Image Generation and ManipulationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Chris Tensmeyer 等AAAI 2022 · 被引用 18 次
- ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-wise Semantic Alignment and GenerationJianan Wang, Guansong Lu, Hang Xu, Zhenguo Li 等CVPR 2022 · 被引用 15 次
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang 等ACM MM 2022 · 被引用 12 次
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 被引用 3 次
它引用的顶会 Paper14
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang 等ICCV 2019 · 被引用 148 次
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm 等ICCV 2019 · 被引用 128 次
- Diverse Image Synthesis From Semantic Layouts via Conditional IMLEKe Li, Tianhao Zhang, Jitendra MalikICCV 2019 · 被引用 102 次
- Sequential Attention GAN for Interactive Image EditingYu Cheng, Zhe Gan, Yitong Li, Jingjing Liu 等ACM MM 2020 · 被引用 74 次
- Describe What to Change: A Text-guided Unsupervised Image-to-image Translation ApproachYahui Liu, Marco De Nadai, Deng Cai, Huayang Li 等ACM MM 2020 · 被引用 58 次
相关 Paper
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningZhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai 等ACM MM 2020 · 被引用 17 次
- UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text EditingLichen Ma, Xiaolong Fu, Gaojing Zhou, Zipeng Guo 等AAAI 2026 · 被引用 1 次
- CoSMo: Content-Style Modulation for Image Retrieval With Text FeedbackSeungmin Lee, Dongwan Kim, Bohyung HanCVPR 2021
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan 等SIGIR 2021 · 被引用 72 次
- FlexIT: Towards Flexible Semantic Image TranslationGuillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk 等CVPR 2022 · 被引用 36 次
