R3CD: Scene Graph to Image Generation with Relation-Aware Compositional Contrastive Control Diffusion
Jinxiu Liu, Qi Liu
Abstract
Image generation tasks have achieved remarkable performance using large-scale diffusion models. However, these models are limited to capturing the abstract relations (viz., interactions excluding positional relations) among multiple entities of complex scene graphs. Two main problems exist: 1) fail to depict more concise and accurate interactions via abstract relations; 2) fail to generate complete entities. To address that, we propose a novel Relation-aware Compositional Contrastive Control Diffusion method, dubbed as R3CD, that leverages large-scale diffusion models to learn abstract interactions from scene graphs. Herein, a scene graph transformer based on node and edge encoding is first designed to perceive both local and global information from input scene graphs, whose embeddings are initialized by a T5 model. Then a joint contrastive loss based on attention maps and denoising steps is developed to control the diffusion model to understand and further generate images, whose spatial structures and interaction features are consistent with a priori relation. Extensive experiments are conducted on two datasets: Visual Genome and COCO-Stuff, and demonstrate that the proposal outperforms existing models both in quantitative and qualitative metrics to generate more realistic and diverse images according to different scene graph specifications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e388a314-06da-428b-9d13-c2b4efc8e025Cited by top-tier papers7
- Scene Graph Disentanglement and Composition for Generalizable Complex Image GenerationYunnan Wang, Ziqiang Li, Wenyao Zhang, Zequn Zhang et al.NeurIPS 2024 · 16 citations
- Controllable 3D Outdoor Scene Generation via Scene GraphsYuheng Liu, Xinke Li, Yuning Zhang, Lu Qi et al.ICCV 2025 · 13 citations
- Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural RectificationGuibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge et al.ICCV 2025 · 1 citation
- Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image GenerationZilyu Ye, Zhiyang Chen, Tiancheng Li, Zemin Huang et al.CVPR 2025
- Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine TranslationAndong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai et al.ACL 2025
Builds on9
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman et al.ICLR 2023 · 361 citations
- Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCYilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum et al.ICML 2023 · 219 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
Related papers
- Hierarchical Image Generation via Transformer-Based Sequential Patch SelectionXiaogang Xu, Ning XuAAAI 2022 · 10 citations
- GraphDreamer: Compositional 3D Scene Synthesis from Scene GraphsGege Gao, Weiyang Liu, Anpei Chen, Andreas Geiger et al.CVPR 2024
- LAW-Diffusion: Complex Scene Generation by Diffusion with LayoutsBinbin Yang, Yi Luo, Ziliang Chen, Guangrun Wang et al.ICCV 2023 · 21 citations
- Exploiting Relationship for Complex-scene Image GenerationTianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang et al.AAAI 2021 · 18 citations
- CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene GraphsGuangyao Zhai, Evin Pinar Örnek, Shun-Cheng Wu, Yan Di et al.NeurIPS 2023 · 76 citations
