Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image Synthesis
Minho Park, Jooyeol Yun, Seunghwan Choi, Jaegul Choo
Abstract
Existing text-to-image generation approaches have set high standards for photorealism and text-image correspondence, largely benefiting from web-scale text-image datasets, which can include up to 5 billion pairs. However, text-to-image generation models trained on domain-specific datasets, such as urban scenes, medical images, and faces, still suffer from low text-image correspondence due to the lack of text-image pairs. Additionally, collecting billions of text-image pairs for a specific domain can be time-consuming and costly. Thus, ensuring high text-image correspondence without relying on web-scale text-image datasets remains a challenging task. In this paper, we present a novel approach for enhancing text-image correspondence by leveraging available semantic layouts. Specifically, we propose a Gaussian-categorical diffusion process that simultaneously generates both images and corresponding layout pairs. Our experiments reveal that we can guide text-to-image generation models to be aware of the semantics of different image regions, by training the model to generate semantic labels for each pixel. We demonstrate that our approach achieves higher text-image correspondence compared to existing text-to-image generation approaches in the Multi-Modal CelebA-HQ and the Cityscapes dataset, where text-image pairs are scarce. Codes are available at https://pmh9960.github.io/research/GCDP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf886bef-d70d-4b42-8e5c-53ab955e93c6Cited by top-tier papers7
- Expressive Text-to-Image Generation with Rich TextSongwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin HuangICCV 2023 · 102 citations
- MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual InvariantChenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang et al.CVPR 2024 · 20 citations
- HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion ModelsWenfeng Song, Xinyu Zhang, Shuai Li, Yang Gao et al.CVPR 2024 · 6 citations
- Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image SynthesisDeepak Sridhar, Abhishek Peri, Rohith Rachala, Nuno VasconcelosNeurIPS 2024 · 5 citations
- Concept-Aware LoRA for Domain-Aligned Segmentation Dataset GenerationMinho Park, Sunghyun Park, Jungsoo Lee, Hyojin Park et al.CVPR 2026 · 1 citation
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Image Synthesis via Semantic CompositionYi Wang, Lu Qi, Ying-Cong Chen, Xiangyu Zhang et al.ICCV 2021 · 72 citations
- Unsupervised Semantic Correspondence Using Stable DiffusionEric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack et al.NeurIPS 2023 · 152 citations
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image GenerationLeigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie et al.ACM MM 2023 · 91 citations
- Semantic Image Analogy with a Conditional Single-Image GANJiacheng Li, Zhiwei Xiong, Dong Liu, Xuejin Chen et al.ACM MM 2020 · 4 citations
- Semantic Palette: Guiding Scene Generation With Class ProportionsGuillaume Le Moing, Tuan-Hung Vu, Himalaya Jain, Patrick Pérez et al.CVPR 2021
