LoCo: Training-Free Layout-to-Image Synthesis with Localized Constraints
Peiang Zhao, Han Li, Ruiyang Jin, S. Kevin Zhou
Abstract
Recent text-to-image diffusion models have achieved remarkable success in generating high-quality images. However, their exclusive reliance on textual prompts falls short in precise control of image compositions. In this paper, we propose LoCo, a training-free approach for layout-to-image synthesis that excels in producing high-quality images aligned with both textual prompts and layout instructions. Specifically, LoCo features a novel Localized Attention Constraint, which utilizes the semantic affinity between pixels in self-attention maps to create precise representations of desired objects, thereby ensuring their accurate placement within designated regions. We further introduce a Padding Token Constraint to leverage the semantic information embedded in previously overlooked padding tokens, improving the consistency between object appearance and layout instructions. Our method seamlessly integrates with existing text-to-image and layout-to-image models, improving their spatial control capabilities and addressing semantic failures seen in prior approaches. Extensive experiments demonstrate the superiority of LoCo, outperforming state-of-the-art training-free layout-to-image methods both qualitatively and quantitatively across multiple benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 245503ae-85a8-4775-bab0-7cbd5144d72bRelated papers
- Zero-Painter: Training-Free Layout Control for Text-to-Image SynthesisMarianna Ohanyan, Hayk Manukyan, Zhangyang Wang, Shant Navasardyan et al.CVPR 2024 · 4 citations
- BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionJinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu et al.ICCV 2023 · 313 citations
- Dense Text-to-Image Generation with Attention ModulationYunji Kim, Jiyoung Lee, Jin-Hwa Kim, Jung-Woo Ha et al.ICCV 2023 · 204 citations
- LAW-Diffusion: Complex Scene Generation by Diffusion with LayoutsBinbin Yang, Yi Luo, Ziliang Chen, Guangrun Wang et al.ICCV 2023 · 21 citations
- Control and Realism: Best of Both Worlds in Layout-to-Image without TrainingBonan Li, Yinhan Hu, Songhua Liu, Xinchao WangICML 2025
