Compositional Text-to-Image Generation with Dense Blob Representations
Weili Nie, Sifei Liu, Morteza Mardani, Chao Liu, Benjamin Eckart, Arash Vahdat
摘要
Existing text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose a scene into visual primitives - denoted as dense blob representations - that contain fine-grained details of the scene while being modular, human-interpretable, and easy-to-construct. Based on blob representations, we develop a blob-grounded text-to-image diffusion model, termed BlobGEN, for compositional generation. Particularly, we introduce a new masked cross-attention module to disentangle the fusion between blob representations and visual features. To leverage the compositionality of large language models (LLMs), we introduce a new in-context learning approach to generate blob representations from text prompts. Our extensive experiments show that BlobGEN achieves superior zero-shot generation quality and better layout-guided controllability on MS-COCO. When augmented by LLMs, our method exhibits superior numerical and spatial correctness on compositional image generation benchmarks. Project page: https://blobgen-2d.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language ModelsJian Ma, Yonglin Deng, Chen Chen, Nanyang Du 等AAAI 2025 · 被引用 28 次
- Kaleido Diffusion: Improving Conditional Diffusion Models with Autoregressive Latent ModelingJiatao Gu, Ying Shen, Shuangfei Zhai, Yizhe Zhang 等NeurIPS 2024 · 被引用 19 次
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose ManipulationZhenyuan Qin, Xincheng Shuai, Henghui DingNeurIPS 2025 · 被引用 11 次
- Hierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language ModelWeilin Sun, Xinran Li, Manyi Li, Kai Xu 等AAAI 2025 · 被引用 7 次
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed IndividualsDavide Lobba, Fulvio Sanguigni, Bin Ren, Marcella Cornia 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video RepresentationsWeixi Feng, Chao Liu, Sifei Liu, William Yang Wang 等CVPR 2025
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan 等ICLR 2024 · 被引用 61 次
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang 等NeurIPS 2024 · 被引用 13 次
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image GenerationLeigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie 等ACM MM 2023 · 被引用 91 次
