NoiseCollage: A Layout-Aware Text-to-Image Diffusion Model Based on Noise Cropping and Merging
Takahiro Shirakawa, Seiichi Uchida
Abstract
Layout-aware text-to-image generation is a task to generate multi-object images that reflect layout conditions in addition to text conditions. The current layout-aware text-to-image diffusion models still have several issues, including mismatches between the text and layout conditions and quality degradation of generated images. This paper proposes a novel layout-aware text-to-image diffusion model called NoiseCollage to tackle these issues. During the denoising process, NoiseCollage independently estimates noises for individual objects and then crops and merges them into a single noise. This operation helps avoid condition mismatches; in other words, it can put the right objects in the right places. Qualitative and quantitative evaluations show that NoiseCollage outperforms several state-of-the-art models. These successful results indicate that the crop-and-merge operation of noises is a reasonable strategy to control image generation. We also show that NoiseCollage can be integrated with ControlNet to use edges, sketches, and pose skeletons as additional conditions. Experimental results show that this integration boosts the layout accuracy of ControlNet. The code is available at https://github.com/univ-esuty/noisecollage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e892359-9d72-495d-bd2b-eb845b55244dCited by top-tier papers11
- GrounDiT: Grounding Diffusion Transformers via Noisy Patch TransplantationYuseung Lee, Taehoon Yoon, Minhyuk SungNeurIPS 2024 · 28 citations
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic ControlDanfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li et al.NeurIPS 2025 · 11 citations
- InstanceAssemble: Layout-Aware Image Generation via Instance Assembling AttentionQiang Xiang, Shuang Sun, Binglei Li, Dejia Song et al.NeurIPS 2025 · 10 citations
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao et al.ICCV 2025 · 7 citations
- Training-Free and Adaptive Sparse Attention for Efficient Long Video GenerationYifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang et al.ICCV 2025 · 6 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- Be Decisive: Noise-Induced Layouts for Multi-Subject GenerationOmer Dahary, Yehonathan Cohen, Or Patashnik, Kfir Aberman et al.SIGGRAPH 2025 · 3 citations
- DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion ModelsHongji Yang, Wencheng Han, Yucheng Zhou, Jianbing ShenICCV 2025 · 4 citations
- SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout ControlJaskirat Singh, Jianming Zhang, Qing Liu, Cameron Smith et al.CVPR 2024 · 7 citations
- DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image GenerationZheng Fang, Lichuan Xiang, Xu Cai, Bing Wang et al.CVPR 2026
- RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion ModelsXinchen Zhang, Ling Yang, Yaqi Cai, Zhaochen Yu et al.NeurIPS 2024 · 22 citations
