Background Layout Generation and Object Knowledge Transfer for Text-to-Image Generation
Zhuowei Chen, Zhendong Mao, Shancheng Fang, Bo Hu
摘要
Text-to-Image generation (T2I) aims to generate realistic and semantically consistent images according to the natural language descriptions. Built upon the recent advances in generative adversarial networks (GANs), existing T2I models have made great process. However, a close inspection of their generated images shows two major limitations: 1) the background (e.g., fence, lake) of the generated image with the complicated, real-world scene tends to be unrealistic; 2) the object (e.g., elephant, zebra) in the generated image often presents highly distorted shape or key parts missing. To address these limitations, we propose a two-stage T2I approach, where the first stage redesigns the text-to-layout process to incorporate the background layout with the existing object layout, the second stage transfers the object knowledge from an existing class-to-image model to the layout-to-image process to improve the object fidelity. Specifically, a transformer-based architecture is introduced as the layout generator to learn the mapping from text to layout of object and background, and a Text-attended Layout-aware feature Normalization (TL-Norm) is proposed to adaptively transfer the object knowledge to the image generation. Benefitting from the background layout and transferred object knowledge, the proposed approach significantly surpasses previous state-of-the-art methods in the image quality metric and achieves superior image-text alignment performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Context-Aware Layout to Image Generation With Enhanced Object AppearanceSen He, Wentong Liao, Michael Ying Yang, Yongxin Yang 等CVPR 2021
- Text to Image Generation with Semantic-Spatial Aware GANWentong Liao, Kai Hu, Michael Ying Yang, Bodo RosenhahnCVPR 2022 · 被引用 169 次
- Text-to-Image Synthesis based on Object-Guided Joint-Decoding TransformerFuxiang Wu, Liu Liu, Fusheng Hao, Fengxiang He 等CVPR 2022 · 被引用 13 次
- GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningYaqi Li, Han Fang, Zerun Feng, Kaijing Ma 等ACM MM 2024 · 被引用 1 次
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten 等SIGGRAPH 2024 · 被引用 57 次
