Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models
Ruichen Wang, Zekang Chen, Chen Chen, Jian Ma, Haonan Lu, Xiaodong Lin
摘要
Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited compositional capabilities, leading to attribute leakage, entity leakage, and missing entities. In this paper, we propose a novel attention mask control strategy based on predicted object boxes to address these issues. In particular, we first train a BoxNet to predict a box for each entity that possesses the attribute specified in the prompt. Then, depending on the predicted boxes, a unique mask control is applied to the cross-and self-attention maps. Our approach produces a more semantically accurate synthesis by constraining the attention regions of each token in the prompt to the image. In addition, the proposed method is straightforward and effective and can be readily integrated into existing cross-attention-based T2I generators. We compare our approach to competing methods and demonstrate that it can faithfully convey the semantics of the original text to the generated content and achieve high availability as a ready-to-use plugin. Please refer to https://github.com/OPPO- Mente-Lab/attention-mask-control. "A black cat and a yellow dog" attribute leakage entity leakage missing entities OURS Figure 1: Example results from Stable Diffusion (first three sets of images) and Our method (last set). Our method aims to address three typical generation defects (attribute leakage, entity leakage, and missing entities) and generate images that are more semantically faithful to the image captions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsLing Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu 等ICML 2024 · 被引用 231 次
- ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise OptimizationLuca Eyring, Shyamgopal Karthik, Karsten Roth, Alexey Dosovitskiy 等NeurIPS 2024 · 被引用 131 次
- Subject-Diffusion: Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuningJian Ma, Junhao Liang, Chen Chen, Haonan LuSIGGRAPH 2024 · 被引用 71 次
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao 等NeurIPS 2024 · 被引用 62 次
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao 等NeurIPS 2024 · 被引用 45 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceZhenhong Sun, Junyan Wang, Zhiyu Tan, Daoyi Dong 等ACM MM 2024 · 被引用 4 次
- ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image GenerationXinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang 等ACM MM 2025
- Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image SynthesisQiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui 等ICCV 2023 · 被引用 55 次
- Easing Concept Bleeding in Diffusion via Entity Localization and AnchoringJiewei Zhang, Song Guo, Peiran Dong, Jie Zhang 等ICML 2024 · 被引用 3 次
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
