Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models
Ruichen Wang, Zekang Chen, Chen Chen, Jian Ma, Haonan Lu, Xiaodong Lin
Abstract
Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited compositional capabilities, leading to attribute leakage, entity leakage, and missing entities. In this paper, we propose a novel attention mask control strategy based on predicted object boxes to address these issues. In particular, we first train a BoxNet to predict a box for each entity that possesses the attribute specified in the prompt. Then, depending on the predicted boxes, a unique mask control is applied to the cross-and self-attention maps. Our approach produces a more semantically accurate synthesis by constraining the attention regions of each token in the prompt to the image. In addition, the proposed method is straightforward and effective and can be readily integrated into existing cross-attention-based T2I generators. We compare our approach to competing methods and demonstrate that it can faithfully convey the semantics of the original text to the generated content and achieve high availability as a ready-to-use plugin. Please refer to https://github.com/OPPO- Mente-Lab/attention-mask-control. "A black cat and a yellow dog" attribute leakage entity leakage missing entities OURS Figure 1: Example results from Stable Diffusion (first three sets of images) and Our method (last set). Our method aims to address three typical generation defects (attribute leakage, entity leakage, and missing entities) and generate images that are more semantically faithful to the image captions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd5fe948-04c8-4a74-b5dd-635f5dafb9b7Cited by top-tier papers34
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsLing Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu et al.ICML 2024 · 231 citations
- ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise OptimizationLuca Eyring, Shyamgopal Karthik, Karsten Roth, Alexey Dosovitskiy et al.NeurIPS 2024 · 131 citations
- Subject-Diffusion: Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuningJian Ma, Junhao Liang, Chen Chen, Haonan LuSIGGRAPH 2024 · 71 citations
- VideoTetris: Towards Compositional Text-to-Video GenerationYe Tian, Ling Yang, Haotian Yang, Yuan Gao et al.NeurIPS 2024 · 62 citations
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao et al.NeurIPS 2024 · 45 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceZhenhong Sun, Junyan Wang, Zhiyu Tan, Daoyi Dong et al.ACM MM 2024 · 4 citations
- ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image GenerationXinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang et al.ACM MM 2025
- Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image SynthesisQiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui et al.ICCV 2023 · 55 citations
- Easing Concept Bleeding in Diffusion via Entity Localization and AnchoringJiewei Zhang, Song Guo, Peiran Dong, Jie Zhang et al.ICML 2024 · 3 citations
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
