R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation
Jiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang, Qingming Huang
Abstract
Recent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout instruction. In this work, we probe into zero-shot grounded T2I generation with diffusion models, that is, generating images corresponding to the input layout information without training auxiliary modules or finetuning diffusion models. We propose a Region and Boundary (R&B) aware cross-attention guidance approach that gradually modulates the attention maps of diffusion model during generative process, and assists the model to synthesize images (1) with high fidelity, (2) highly compatible with textual input, and (3) interpreting layout instructions accurately. Specifically, we leverage the discrete sampling to bridge the gap between consecutive attention maps and discrete layout constraints, and design a region-aware loss to refine the generative layout during diffusion process. We further propose a boundary-aware loss to strengthen object discriminability within the corresponding regions. Experimental results show that our method outperforms existing state-of-the-art zero-shot grounded T2I generation methods by a large margin both qualitatively and quantitatively on several benchmarks. Project page: https://sagileo.github.io/Region-and-Boundary .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5958168e-a6ff-490c-b7ed-c213e435df1bCited by top-tier papers21
- Grounded Text-to-Image Synthesis with Attention RefocusingQuynh Phung, Songwei Ge, Jia-Bin HuangCVPR 2024 · 59 citations
- Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without GuidanceKuan Heng Lin, Sicheng Mo, Ben Klingher, Fangzhou Mu et al.NeurIPS 2024 · 51 citations
- GrounDiT: Grounding Diffusion Transformers via Noisy Patch TransplantationYuseung Lee, Taehoon Yoon, Minhyuk SungNeurIPS 2024 · 28 citations
- NoiseCollage: A Layout-Aware Text-to-Image Diffusion Model Based on Noise Cropping and MergingTakahiro Shirakawa, Seiichi UchidaCVPR 2024 · 19 citations
- PLACE: Adaptive Layout-Semantic Fusion for Semantic Image SynthesisZhengyao Lv, Yuxiang Wei, Wangmeng Zuo, Kwan-Yee K. WongCVPR 2024 · 14 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- Zero-shot spatial layout conditioning for text-to-image diffusion modelsGuillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière et al.ICCV 2023 · 82 citations
- SpotActor: Training-Free Layout-Controlled Consistent Image GenerationJiahao Wang, Caixia Yan, Weizhan Zhang, Haonan Lin et al.AAAI 2025 · 13 citations
- Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image SynthesisQiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui et al.ICCV 2023 · 55 citations
- SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image GenerationChengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai et al.AAAI 2024 · 30 citations
