Sequential Attention GAN for Interactive Image Editing
Yu Cheng, Zhe Gan, Yitong Li, Jingjing Liu, Jianfeng Gao
Abstract
Most existing text-to-image synthesis tasks are static single-turn generation, based on pre-defined textual descriptions of images. To explore more practical and interactive real-life applications, we introduce a new task - Interactive Image Editing, where users can guide an agent to edit images via multi-turn textual commands on-the-fly. In each session, the agent takes a natural language description from the user as the input, and modifies the image generated in previous turn to a new design, following the user description. The main challenges in this sequential and interactive image generation task are two-fold: 1) contextual consistency between a generated image and the provided textual description; 2) step-by-step region-level modification to maintain visual consistency across the generated image sequence in each session. To address these challenges, we propose a novel Sequential Attention Generative Adversarial Network (SeqAttnGAN), which applies a neural state tracker to encode the previous image and the textual description in each turn of the sequence, and uses a GAN framework to generate a modified version of the image that is consistent with the preceding images and coherent with the description. To achieve better region-specific refinement, we also introduce a sequential attention mechanism into the model. To benchmark on the new task, we introduce two new datasets, Zap-Seq and DeepFashion-Seq, which contain multi-turn sessions with image-description sequences in the fashion domain. Experiments on both datasets show that the proposed SeqAttnGAN model outperforms state-of-the-art approaches on the interactive image editing task across all evaluation metrics including visual quality, image sequence coherence and text-image consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28b38294-29cd-42f5-873c-cec8c315000aCited by top-tier papers20
- Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Literature ReviewYang Shi, Tian Gao, Xiaohan Jiao, Nan CaoCSCW 2023 · 170 citations
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy et al.ICCV 2021 · 162 citations
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 76 citations
- Language-based Photo Color Adjustment for Graphic DesignsZhenwei Wang, Nanxuan Zhao, Gerhard P. Hancke, Rynson W. H. LauSIGGRAPH 2023 · 49 citations
- PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape RenderingRong Huang, Haichuan Lin, Chuanzhang Chen, Kang Zhang et al.CHI 2024 · 45 citations
Builds on2
Related papers
- Semantics-Enhanced Adversarial Nets for Text-to-Image SynthesisHongchen Tan, Xiuping Liu, Xin Li, Yi Zhang et al.ICCV 2019 · 80 citations
- Fashion Editing With Adversarial Parsing LearningHaoye Dong, Xiaodan Liang, Yixuan Zhang, Xujie Zhang et al.CVPR 2020
- TiGAN: Text-Based Interactive Image Generation and ManipulationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Chris Tensmeyer et al.AAAI 2022 · 18 citations
- FACT: Fused Attention for Clothing Transfer with Generative Adversarial NetworksYicheng Zhang, Lei Li, Li Song, Rong Xie et al.AAAI 2020 · 11 citations
- LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningGaoxiang Cong, Liang Li, Zhenhuan Liu, Yunbin Tu et al.ACM MM 2022 · 11 citations
