R3CD: Scene Graph to Image Generation with Relation-Aware Compositional Contrastive Control Diffusion
Jinxiu Liu, Qi Liu
摘要
Image generation tasks have achieved remarkable performance using large-scale diffusion models. However, these models are limited to capturing the abstract relations (viz., interactions excluding positional relations) among multiple entities of complex scene graphs. Two main problems exist: 1) fail to depict more concise and accurate interactions via abstract relations; 2) fail to generate complete entities. To address that, we propose a novel Relation-aware Compositional Contrastive Control Diffusion method, dubbed as R3CD, that leverages large-scale diffusion models to learn abstract interactions from scene graphs. Herein, a scene graph transformer based on node and edge encoding is first designed to perceive both local and global information from input scene graphs, whose embeddings are initialized by a T5 model. Then a joint contrastive loss based on attention maps and denoising steps is developed to control the diffusion model to understand and further generate images, whose spatial structures and interaction features are consistent with a priori relation. Extensive experiments are conducted on two datasets: Visual Genome and COCO-Stuff, and demonstrate that the proposal outperforms existing models both in quantitative and qualitative metrics to generate more realistic and diverse images according to different scene graph specifications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Scene Graph Disentanglement and Composition for Generalizable Complex Image GenerationYunnan Wang, Ziqiang Li, Wenyao Zhang, Zequn Zhang 等NeurIPS 2024 · 被引用 16 次
- Controllable 3D Outdoor Scene Generation via Scene GraphsYuheng Liu, Xinke Li, Yuning Zhang, Lu Qi 等ICCV 2025 · 被引用 13 次
- Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural RectificationGuibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge 等ICCV 2025 · 被引用 1 次
- Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image GenerationZilyu Ye, Zhiyang Chen, Tiancheng Li, Zemin Huang 等CVPR 2025
- Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine TranslationAndong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai 等ACL 2025
它引用的顶会 Paper9
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
- Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCYilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum 等ICML 2023 · 被引用 219 次
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 被引用 190 次
相关 Paper
- Hierarchical Image Generation via Transformer-Based Sequential Patch SelectionXiaogang Xu, Ning XuAAAI 2022 · 被引用 10 次
- GraphDreamer: Compositional 3D Scene Synthesis from Scene GraphsGege Gao, Weiyang Liu, Anpei Chen, Andreas Geiger 等CVPR 2024
- LAW-Diffusion: Complex Scene Generation by Diffusion with LayoutsBinbin Yang, Yi Luo, Ziliang Chen, Guangrun Wang 等ICCV 2023 · 被引用 21 次
- Exploiting Relationship for Complex-scene Image GenerationTianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang 等AAAI 2021 · 被引用 18 次
- CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene GraphsGuangyao Zhai, Evin Pinar Örnek, Shun-Cheng Wu, Yan Di 等NeurIPS 2023 · 被引用 76 次
