Grounded Text-to-Image Synthesis with Attention Refocusing
Quynh Phung, Songwei Ge, Jia-Bin Huang
Abstract
Driven by the scalable diffusion models trained on large-scale datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail to pre-cisely follow the text prompt involving multiple objects, attributes, or spatial compositions. In this paper, we reveal the potential causes in the diffusion model's cross-attention and self-attention layers. We propose two novel losses to refocus attention maps according to a given spatial layout during sampling. Creating the layouts manually requires additional effort and can be tedious. Therefore, we explore using large language models (LLM) to produce these lay-outs for our method. We conduct extensive experiments on the DrawBench, HRS, and TIFA benchmarks to evaluate our proposed method. We show that our proposed attention re-focusing effectively improves the controllability of existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bdba204-9061-43cf-a5cd-8b1537739df3Cited by top-tier papers80
- BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionJinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu et al.ICCV 2023 · 313 citations
- Dense Text-to-Image Generation with Attention ModulationYunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha et al.ICCV 2023 · 204 citations
- LLM-grounded Video Diffusion ModelsLong Lian, Baifeng Shi, Adam Yala, Trevor Darrell et al.ICLR 2024 · 87 citations
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept MatchingDongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang et al.NeurIPS 2024 · 75 citations
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan et al.ICLR 2024 · 61 citations
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image GenerationLeigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie et al.ACM MM 2023 · 91 citations
- Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image SynthesisQiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui et al.ICCV 2023 · 55 citations
- MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic ExpansionFei Peng, Junqiang Wu, Yan Li, Tingting Gao et al.ICCV 2025 · 1 citation
- CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepZheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang et al.NeurIPS 2025 · 9 citations
