ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-wise Semantic Alignment and Generation
Jianan Wang, Guansong Lu, Hang Xu, Zhenguo Li, Chunjing Xu, Yanwei Fu
摘要
Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical application. In this work, we study a novel task on text-guided image manipulation on the entity level in the real world. The task imposes three basic requirements, (1) to edit the entity consistent with the text descriptions, (2) to preserve the text-irrelevant regions, and (3) to merge the manipulated entity into the image naturally. To this end, we propose a new transformer-based framework based on the two-stage image synthesis method, namely ManiTrans, which can not only edit the appearance of entities but also generate new entities corresponding to the text guidance. Our framework incorporates a semantic alignment module to locate the image regions to be manipulated, and a semantic loss to help align the relationship between the vision and language. We conduct extensive experiments on the real datasets, CUB, Oxford, and COCO datasets to verify that our method can distinguish the relevant and irrelevant regions and achieve more precise and flexible manipulation compared with baseline methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 被引用 102 次
- High Quality Entity SegmentationLu Qi, Jason Kuen, Tiancheng Shen, Jiuxiang Gu 等ICCV 2023 · 被引用 91 次
- Towards Efficient Diffusion-Based Image Editing with Instant Attention MasksSiyu Zou, Jiji Tang, Yiyi Zhou, Jing He 等AAAI 2024 · 被引用 24 次
- DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksMing Tao, Bing-Kun Bao, Hao Tang, Fei Wu 等AAAI 2023 · 被引用 19 次
- Multi-Region Text-Driven Manipulation of Diffusion ImageryYiming Li, Peng Zhou, Jun Sun, Yi XuAAAI 2024 · 被引用 4 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
相关 Paper
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
- Text as Neural Operator: Image Manipulation by Text InstructionTianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang 等ACM MM 2021 · 被引用 28 次
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 被引用 3 次
- Cycle-Consistent Inverse GAN for Text-to-Image SynthesisHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2021 · 被引用 47 次
- Background Layout Generation and Object Knowledge Transfer for Text-to-Image GenerationZhuowei Chen, Zhendong Mao, Shancheng Fang, Bo HuACM MM 2022 · 被引用 6 次
