Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
Yunnan Wang, Ziqiang Li, Wenyao Zhang, Zequn Zhang, Baao Xie, Xihui Liu, Wenjun Zeng, Xin Jin
摘要
There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an"isolated"image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic ReasoningYuhui Zeng, Haoxiang Wu, Wenjie Nie, Guangyao Chen 等ICCV 2025 · 被引用 2 次
- SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic AnnotationsYunnan Wang, Kecheng Zheng, Jianyuan Wang, Minghao Chen 等CVPR 2026 · 被引用 1 次
- A-Bench: Are LMMs Masters at Evaluating AI-generated Images?Zicheng Zhang, Haoning Wu, Chunyi Li, Yingjie Zhou 等ICLR 2025
- Inversion-DPO: Precise and Efficient Post-Training for Diffusion ModelsZejian Li, Yize Li, Chenye Meng, Zhongni Liu 等ACM MM 2025
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Exploiting Relationship for Complex-scene Image GenerationTianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang 等AAAI 2021 · 被引用 18 次
- VarScene: A Deep Generative Model for Realistic Scene Graph SynthesisTathagat Verma, Abir De, Yateesh Agrawal, Vishwa Vinay 等ICML 2022 · 被引用 11 次
- SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB SequencesSeok-Young Kim, Dooyoung Kim, Woojin Cho, Hail Song 等IEEE VR 2026 · 被引用 1 次
- End-to-End Optimization of Scene LayoutAndrew Luo, Zhoutong Zhang, Jiajun Wu, Joshua B. TenenbaumCVPR 2020
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang 等AAAI 2025 · 被引用 1 次
