Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
Yunnan Wang, Ziqiang Li, Wenyao Zhang, Zequn Zhang, Baao Xie, Xihui Liu, Wenjun Zeng, Xin Jin
Abstract
There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an"isolated"image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic ReasoningYuhui Zeng, Haoxiang Wu, Wenjie Nie, Guangyao Chen et al.ICCV 2025 · 2 citations
- SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic AnnotationsYunnan Wang, Kecheng Zheng, Jianyuan Wang, Minghao Chen et al.CVPR 2026 · 1 citation
- A-Bench: Are LMMs Masters at Evaluating AI-generated Images?Zicheng Zhang, Haoning Wu, Chunyi Li, Yingjie Zhou et al.ICLR 2025
- Inversion-DPO: Precise and Efficient Post-Training for Diffusion ModelsZejian Li, Yize Li, Chenye Meng, Zhongni Liu et al.ACM MM 2025
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Exploiting Relationship for Complex-scene Image GenerationTianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang et al.AAAI 2021 · 18 citations
- VarScene: A Deep Generative Model for Realistic Scene Graph SynthesisTathagat Verma, Abir De, Yateesh Agrawal, Vishwa Vinay et al.ICML 2022 · 11 citations
- SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB SequencesSeok-Young Kim, Dooyoung Kim, Woojin Cho, Hail Song et al.IEEE VR 2026 · 1 citation
- End-to-End Optimization of Scene LayoutAndrew Luo, Zhoutong Zhang, Jiajun Wu, Joshua B. TenenbaumCVPR 2020
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang et al.AAAI 2025 · 1 citation
