Lune

ACM MM2025Top-tier venue

Text-to-Image Generation with Multi-modal Knowledge Graph Construction and Retrieval

Jiawei Meng, Zhengmao Yang, Zhiqiang Liu, Shaokai Chen, Zhizhen Liu, Wen Zhang, Huajun Chen

2025Year

Abstract

Current Text-to-Image (T2I) generation methods struggle to accurately create images with complex object relationships and scene compositions. To overcome these challenges, we propose KAIG, a novel text-to-image generative model that integrates a knowledge graph into the image generation process. Unlike traditional models, KAIG uses structured knowledge to enhance the retrieval of relevant information, enabling the generation of high-quality, contextually rich, and semantically consistent images from multi-modal inputs. We introduce a two-stage training strategy: first, condition adapters are trained to align multi-modal inputs, followed by fine-tuning the entire diffusion model. This approach ensures precise alignment between retrieved conditions and the image generation process, leading to an efficient and scalable pipeline. Our experiments on two popular datasets, MS-COCO and CUB-200-2011, show that KAIG consistently outperforms existing methods in both image quality and consistency. Notably, KAIG can seamlessly integrate with any pre-trained diffusion model, requiring minimal additional training while achieving superior results. Ultimately, KAIG demonstrates strong potential for addressing key limitations in current T2I models and advancing the field of image synthesis.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ef51f474-7c17-44f4-9df9-e72053e49200

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines