StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis
Zhiheng Li, Martin Renqiang Min, Kai Li, Chenliang Xu
摘要
Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lacking compositionality could have severe implications for robustness and fairness, e.g., inability to synthesize the face images of underrepresented demographic groups. In this paper, we introduce a new framework, StyleT2I, to improve the compositionality of text-to-image synthesis. Specifically, we propose a CLIP-guided Contrastive Loss to better distinguish different compositions among different sentences. To further improve the compositionality, we design a novel Semantic Matching Loss and a Spatial Constraint to identify attributes' latent directions for intended spatial region manipulations, leading to better disentangled latent representations of attributes. Based on the identified latent directions of attributes, we propose Compositional Attribute Adjustment to adjust the latent code, resulting in better compositionality of image synthesis. In addition, we leverage the <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> -norm regularization of identified latent directions (norm penalty) to strike a nice balance between image-text alignment and image fidelity. In the experiments, we devise a new dataset split and an evaluation metric to evaluate the compositionality of text-to-image synthesis models. The results show that StyleT2I outperforms previous approaches in terms of the consistency between the input text and synthesized images and achieves higher fidelity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang 等NeurIPS 2023 · 被引用 119 次
- Effective pruning of web-scale datasets based on complexity of concept clustersAmro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel 等ICLR 2024 · 被引用 30 次
- Are Diffusion Models Vision-And-Language Reasoners?Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Chris Pal 等NeurIPS 2023 · 被引用 21 次
- Attention Calibration for Disentangled Text-to-Image PersonalizationYanbing Zhang, Mengping Yang, Qin Zhou, Zhe WangCVPR 2024 · 被引用 17 次
- SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color EditingJing Shi, Ning Xu, Haitian Zheng, Alex Smith 等CVPR 2022 · 被引用 15 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
相关 Paper
- Towards Counterfactual Image Manipulation via CLIPYingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang 等ACM MM 2022 · 被引用 33 次
- Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic DataHaoxin Li, Boyang LiCVPR 2025
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 被引用 1 次
- SAT3D: Image-driven Semantic Attribute Transfer in 3DZhijun Zhai, Zengmao Wang, Xiaoxiao Long, Kaixuan Zhou 等ACM MM 2024
- Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image ManipulationXiwen Wei, Zhen Xu, Cheng Liu, Si Wu 等CVPR 2023
