Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, Graham W. Taylor
Abstract
Conditional text-to-image generation is an active area of research, with many possible applications. Existing research has primarily focused on generating a single image from available conditioning information in one step. One practical extension beyond one-step generation is a system that generates an image iteratively, conditioned on ongoing linguistic input or feedback. This is significantly more challenging than one-step generation tasks, as such a system must understand the contents of its generated images with respect to the feedback history, the current feedback, as well as the interactions among concepts present in the feedback history. In this work, we present a recurrent image generation model which takes into account both the generated output up to the current step as well as all past instructions for generation. We show that our model is able to generate the background, add new objects, and apply simple transformations to existing objects. We believe our approach is an important step toward interactive generation. Code and data is available at: https://www.microsoft.com/en-us/research/ project/generative-neural-visual-artist-geneva/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00dce0e1-757f-4872-a0ce-b01acbfdd40eCited by top-tier papers32
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- Object-Centric Image Generation from LayoutsTristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm et al.AAAI 2021 · 107 citations
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 76 citations
Related papers
- Text as Neural Operator: Image Manipulation by Text InstructionTianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang et al.ACM MM 2021 · 28 citations
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningZhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai et al.ACM MM 2020 · 17 citations
- Draft-and-Revise: Effective Image Generation with Contextual RQ-TransformerDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.NeurIPS 2022 · 36 citations
- LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningGaoxiang Cong, Liang Li, Zhenhuan Liu, Yunbin Tu et al.ACM MM 2022 · 11 citations
- TiGAN: Text-Based Interactive Image Generation and ManipulationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Chris Tensmeyer et al.AAAI 2022 · 18 citations
