EGGen: Image Generation with Multi-entity Prior Learning through Entity Guidance
Zhenhong Sun, Junyan Wang, Zhiyu Tan, Daoyi Dong, Hailan Ma, Hao Li, Dong Gong
摘要
Diffusion models have shown remarkable prowess in text-to-image synthesis and editing, yet they often stumble when tasked with interpreting complex prompts that describe multiple entities with specific attributes and interrelations. The generated images often contain inconsistent multi-entity representation (IMR), reflected as inaccurate presentations of the multiple entities and their attributes. Although providing spatial layout guidance improves the multientity generation quality in existing works, it is still challenging to handle the leakage attributes and avoid unnatural characteristics. To address the IMR challenge, we first conduct in-depth analyses of the diffusion process and attention operation, revealing that the IMR challenges largely stem from the process of cross-attention mechanisms. According to the analyses, we introduce the entity guidance generation mechanism, which maintains the integrity of the original diffusion model parameters by integrating plugin networks. Our work advances the stable diffusion model by segmenting comprehensive prompts into distinct entity-specific prompts with bounding boxes, enabling a transition from multientity to single-entity generation in cross-attention layers. More importantly, we introduce entity-centric cross-attention layers that focus on individual entities to preserve their uniqueness and accuracy, alongside global entity alignment layers that refine crossattention maps using multi-entity priors for precise positioning and attribute accuracy. Additionally, a linear attenuation module is integrated to progressively reduce the influence of these layers during inference, preventing oversaturation and preserving generation fidelity. Our comprehensive experiments demonstrate that this entity guidance generation enhances existing text-to-image models in generating detailed, multi-entity images. Code is available at https://github.com/chaos-sun/eggen.git.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang 等CVPR 2026 · 被引用 11 次
- DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingXin Xie, Dong GongCVPR 2025
- VODiff: Controlling Object Visibility Order in Text-to-Image GenerationDong Liang, Jinyuan Jia, Yuhao Liu, Zhanghan Ke 等CVPR 2025
- Prototype-Based Image Prompting for Weakly Supervised Histopathological Image SegmentationQingchen Tang, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
相关 Paper
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
- Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion ModelsRuichen Wang, Zekang Chen, Chen Chen, Jian Ma 等AAAI 2024 · 被引用 97 次
- CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion ModelsTuna Han Salih Meral, Enis Simsar, Federico Tombari, Pinar YanardagCVPR 2024 · 被引用 11 次
- Easing Concept Bleeding in Diffusion via Entity Localization and AnchoringJiewei Zhang, Song Guo, Peiran Dong, Jie Zhang 等ICML 2024 · 被引用 3 次
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance GenerationRuihang Xu, Dewei Zhou, Fan Ma, Yi YangICLR 2026 · 被引用 19 次
