LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts
Binbin Yang, Yi Luo, Ziliang Chen, Guangrun Wang, Xiaodan Liang, Liang Lin
摘要
Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a text is often too abstract to properly specify all the spatial properties of an image, e.g., the layout configuration of a scene, leading to the sub-optimal results of complex scene generation. In this paper, we achieve accurate complex scene generation by proposing a semantically controllable Layout-AWare diffusion model, termed LAW-Diffusion. Distinct from the previous Layout-to-Image generation (L2I) methods that primarily explore category-aware relationships, LAW-Diffusion introduces a spatial dependency parser to encode the location-aware semantic coherence across objects as a layout embedding and produces a scene with perceptually harmonious object styles and contextual relations. To be specific, we delicately instantiate each object’s regional semantics as an object region map and leverage a location-aware cross-object attention module to capture the spatial dependencies among those disentangled representations. We further propose an adaptive guidance schedule for our layout guidance to mitigate the trade-off between the regional semantic alignment and the texture fidelity of generated objects. Moreover, LAW-Diffusion allows for instance reconfiguration while maintaining the other regions in a synthesized image by introducing a layout-aware latent grafting mechanism to recompose its local regional semantics. To better verify the plausibility of generated scenes, we propose a new evaluation metric for the L2I task, dubbed Scene Relation Score (SRS) to measure how the images preserve the rational and harmonious relations among contextual objects. Comprehensive experiments on COCO-Stuff and Visual-Genome demonstrate that our LAW-Diffusion yields the state-of-the-art generative performance, especially with coherent object relations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic ControlDanfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li 等NeurIPS 2025 · 被引用 11 次
- InstanceAssemble: Layout-Aware Image Generation via Instance Assembling AttentionQiang Xiang, Shuang Sun, Binglei Li, Dejia Song 等NeurIPS 2025 · 被引用 10 次
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao 等ICCV 2025 · 被引用 7 次
- DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous DrivingLiuhan Yin, Runkun Ju, Guodong Guo, Erkang ChengAAAI 2026 · 被引用 4 次
- Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior AestheticsKeming Wu, Junwen Chen, Zhanhao Liang, Yinuo Wang 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image GenerationChengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai 等AAAI 2024 · 被引用 30 次
- LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image GenerationGuangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi 等CVPR 2023
- Context-Aware Layout to Image Generation With Enhanced Object AppearanceSen He, Wentong Liao, Michael Ying Yang, Yongxin Yang 等CVPR 2021
- Exploiting Relationship for Complex-scene Image GenerationTianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang 等AAAI 2021 · 被引用 18 次
- CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepZheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang 等NeurIPS 2025 · 被引用 9 次
