Sequential Attention GAN for Interactive Image Editing
Yu Cheng, Zhe Gan, Yitong Li, Jingjing Liu, Jianfeng Gao
摘要
Most existing text-to-image synthesis tasks are static single-turn generation, based on pre-defined textual descriptions of images. To explore more practical and interactive real-life applications, we introduce a new task - Interactive Image Editing, where users can guide an agent to edit images via multi-turn textual commands on-the-fly. In each session, the agent takes a natural language description from the user as the input, and modifies the image generated in previous turn to a new design, following the user description. The main challenges in this sequential and interactive image generation task are two-fold: 1) contextual consistency between a generated image and the provided textual description; 2) step-by-step region-level modification to maintain visual consistency across the generated image sequence in each session. To address these challenges, we propose a novel Sequential Attention Generative Adversarial Network (SeqAttnGAN), which applies a neural state tracker to encode the previous image and the textual description in each turn of the sequence, and uses a GAN framework to generate a modified version of the image that is consistent with the preceding images and coherent with the description. To achieve better region-specific refinement, we also introduce a sequential attention mechanism into the model. To benchmark on the new task, we introduce two new datasets, Zap-Seq and DeepFashion-Seq, which contain multi-turn sessions with image-description sequences in the fashion domain. Experiments on both datasets show that the proposed SeqAttnGAN model outperforms state-of-the-art approaches on the interactive image editing task across all evaluation metrics including visual quality, image sequence coherence and text-image consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Literature ReviewYang Shi, Tian Gao, Xiaohan Jiao, Nan CaoCSCW 2023 · 被引用 170 次
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy 等ICCV 2021 · 被引用 162 次
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 被引用 76 次
- Language-based Photo Color Adjustment for Graphic DesignsZhenwei Wang, Nanxuan Zhao, Gerhard P. Hancke, Rynson W. H. LauSIGGRAPH 2023 · 被引用 49 次
- PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape RenderingRong Huang, Haichuan Lin, Chuanzhang Chen, Kang Zhang 等CHI 2024 · 被引用 45 次
它引用的顶会 Paper2
相关 Paper
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang 等ICCV 2019 · 被引用 80 次
- Fashion Editing With Adversarial Parsing LearningHaoye Dong, Xiaodan Liang, Yixuan Zhang, Xujie Zhang 等CVPR 2020
- TiGAN: Text-Based Interactive Image Generation and ManipulationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Chris Tensmeyer 等AAAI 2022 · 被引用 18 次
- FACT: Fused Attention for Clothing Transfer with Generative Adversarial NetworksYicheng Zhang, Lei Li, Li Song, Rong Xie 等AAAI 2020 · 被引用 11 次
- LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningGaoxiang Cong, Liang Li, Zhenhuan Liu, Yunbin Tu 等ACM MM 2022 · 被引用 11 次
