Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification
Guibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge, Chaozhe Zhang, Xin Tao, Di Zhang, Pengfei Wan, Guangyong Chen, Yijun Li, Ying-Cong Chen
Abstract
Recent advancements in text-to-image generation have been propelled by the development of diffusion models and multimodality learning. However, since text is typically represented sequentially in these models, it often falls short in providing accurate contextualization and structural control. So the generated images do not consistently align with human expectations, especially in complex scenarios involving multiple objects and relationships. In this paper, we introduce the Scene Graph Adapter (SG-Adapter), leveraging the structured representation of scene graphs to rectify inaccuracies in the original text embeddings. The SG-Adapter's explicit, non-fully connected graph representation significantly improves upon the causal connections commonly used in transformer-based text models. In causal connections, each token can attend to all previous tokens, which may result in attribute leakage. On the other hand, we also curated a highly clean, multi-relational scene graph-image paired dataset MultiRels to address the challenges posed by low-quality annotated datasets like Visual Genome [13]. Furthermore, we design three metrics derived from GPT-4V [1] to effectively and thoroughly measure the correspondence between images and scene graphs. Both qualitative and quantitative results validate the efficacy of our approach in controlling the correspondence in multiple relationships.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01b77920-135c-4ccf-9258-bdcab613b43dBuilds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- Text-to-Image Generation with Multi-modal Knowledge Graph Construction and RetrievalJiawei Meng, Zhengmao Yang, Zhiqiang Liu, Shaokai Chen et al.ACM MM 2025
- Scene Graph Disentanglement and Composition for Generalizable Complex Image GenerationYunnan Wang, Ziqiang Li, Wenyao Zhang, Zequn Zhang et al.NeurIPS 2024 · 16 citations
- Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph GenerationYiming Li, Sisi You, Bing-Kun BaoCVPR 2026
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang et al.AAAI 2025 · 1 citation
- R3CD: Scene Graph to Image Generation with Relation-Aware Compositional Contrastive Control DiffusionJinxiu Liu, Qi LiuAAAI 2024 · 21 citations
