Dual Adversarial Inference for Text-to-Image Synthesis
Qicheng Lao, Mohammad Havaei, Ahmad Pesaranghader, Francis Dutil, Lisa Di-Jorio, Thomas Fevens
摘要
Synthesizing images from a given text description involves engaging two types of information: the content, which includes information explicitly described in the text (e.g., color, composition, etc.), and the style, which is usually not well described in the text (e.g., location, quantity, size, etc.). However, in previous works, it is typically treated as a process of generating images only from the content, i.e., without considering learning meaningful style representations. In this paper, we aim to learn two variables that are disentangled in the latent space, representing content and style respectively. We achieve this by augmenting current text-to-image synthesis frameworks with a dual adversarial inference mechanism. Through extensive experiments, we show that our model learns, in an unsupervised manner, style representations corresponding to certain meaningful information present in the image that are not well described in the text. The new framework also improves the quality of synthesized images when evaluated on Oxford-102, CUB and COCO datasets. Content sources from text descriptions This flower has petals that are pink and has yellow stamen. This flower has petals that are yellow and has dark lines. This flower has white petals as well as a pedicel.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen 等CVPR 2022 · 被引用 607 次
- DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Sha, Zheng Li, Ning Yu, Yang ZhangCCS 2023 · 被引用 123 次
- ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion DesignXujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie 等ACM MM 2022 · 被引用 27 次
- DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic AlignmentXujie Zhang, Binbin Yang, Michael C. Kampffmeyer, Wenqing Zhang 等ICCV 2023 · 被引用 23 次
- ZeroFake: Zero-Shot Detection of Fake Images Generated and Edited by Text-to-Image Generation ModelsZeyang Sha, Yicong Tan, Mingjie Li, Michael Backes 等CCS 2024 · 被引用 8 次
相关 Paper
- Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingBoqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin WangCVPR 2024 · 被引用 9 次
- Unsupervised Compositional Concepts Discovery with Text-to-Image Generative ModelsNan Liu, Yilun Du, Shuang Li, Joshua B. Tenenbaum 等ICCV 2023 · 被引用 40 次
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon 等ACL 2020 · 被引用 66 次
- Disentangled Learning with Synthetic Parallel Data for Text Style TransferJingxuan Han, Quan Wang, Zikang Guo, Benfeng Xu 等ACL 2024 · 被引用 4 次
- Adversarial Disentanglement with Grouped ObservationsJózsef NémethAAAI 2020 · 被引用 8 次
