Imagine That! Abstract-to-Intricate Text-to-Image Synthesis with Scene Graph Hallucination Diffusion
Shengqiong Wu, Hao Fei, Hanwang Zhang, Tat-Seng Chua
Abstract
In this work, we investigate the task of text-to-image (T2I) synthesis under the abstract-to-intricate setting, i.e., generating intricate visual content from simple abstract text prompts. Inspired by human imagination intuition, we propose a novel scene-graph hallucination (SGH) mechanism for effective abstract-to-intricate T2I synthesis. SGH carries out scene hallucination by expanding the initial scene graph (SG) of the input prompt with more feasible specific scene structures, in which the structured semantic representation of SG ensures high controllability of the intrinsic scene imagination. To approach the T2I synthesis, we deliberately build an SG-based hallucination diffusion system. First, we implement the SGH module based on the discrete diffusion technique, which evolves the SG structure by iteratively adding new scene elements. Then, we utilize another continuous-state diffusion model as the T2I synthesizer, where the overt image-generating process is navigated by the underlying semantic scene structure induced from the SGH module. On the benchmark COCO dataset, our system outperforms the existing best-performing T2I model by a significant margin, especially improving on the abstract-to-intricate T2I generation. Further in-depth analyses reveal how our methods advance. 2 * Hao Fei is the corresponding author. 2 Code is available at https://github.com/ChocoWu/T2I-Salad 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang et al.ICML 2024 · 182 citations
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, EditingHao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua et al.NeurIPS 2024 · 100 citations
- Inkspire: Supporting Design Exploration with Generative AI through Analogical SketchingDavid Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro, Aniket Kittur et al.CHI 2025 · 56 citations
- Combating Multimodal LLM Hallucination via Bottom-Up Holistic ReasoningShengqiong Wu, Hao Fei, Liangming Pan, William Yang Wang et al.AAAI 2025 · 24 citations
- PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment AnalysisMeng Luo, Hao Fei, Bobo Li, Shengqiong Wu et al.ACM MM 2024 · 23 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- R3CD: Scene Graph to Image Generation with Relation-Aware Compositional Contrastive Control DiffusionJinxiu Liu, Qi LiuAAAI 2024 · 21 citations
- Text-to-Image Generation with Multi-modal Knowledge Graph Construction and RetrievalJiawei Meng, Zhengmao Yang, Zhiqiang Liu, Shaokai Chen et al.ACM MM 2025
- GraphDreamer: Compositional 3D Scene Synthesis from Scene GraphsGege Gao, Weiyang Liu, Anpei Chen, Andreas Geiger et al.CVPR 2024
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioningXu Zhang, Jin Yuan, Hanwang Zhang, Guojin Zhong et al.AAAI 2025 · 2 citations
- Language-driven Scene Synthesis using Multi-conditional Diffusion ModelVuong Dinh An, Minh Nhat Vu, Toan Nguyen, Baoru Huang et al.NeurIPS 2023 · 14 citations
