ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
Ruihang Xu, Dewei Zhou, Fan Ma, Yi Yang
Abstract
Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserving the identity of multiple distinct subjects. To address these limitations, we introduce ContextGen, a novel Diffusion Transformer framework for multi-instance generation that is guided by both layout and reference images. Our approach integrates two key technical contributions: a Contextual Layout Anchoring (CLA) mechanism that incorporates the composite layout image into the generation context to robustly anchor the objects in their desired positions, and Identity Consistency Attention (ICA), an innovative attention mechanism that leverages contextual reference images to ensure the identity consistency of multiple instances. To address the absence of a large-scale, high-quality dataset for this task, we introduce IMIG-100K, the first dataset to provide detailed layout and identity annotations specifically designed for Multi-Instance Generation. Extensive experiments demonstrate that ContextGen sets a new state-of-the-art, outperforming existing methods especially in layout control and identity fidelity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e76ba69-244d-4852-8e8d-34165d2f953fCited by top-tier papers5
- Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion TransformersRuidong Chen, Yancheng Bai, Xuanpu Zhang, Jianhao Zeng et al.CVPR 2026 · 9 citations
- ConsistCompose: Unified Multimodal Layout Control for Image CompositionXuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai et al.CVPR 2026 · 5 citations
- Are Image-to-Video Models Good Zero-Shot Image Editors?Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi YangCVPR 2026 · 4 citations
- FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured ScriptsYou Li, Dewei Zhou, Fan Ma, Fu Li et al.CVPR 2026 · 2 citations
- DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video GenerationYang-yang Li, Wu Liu, Jie Li, Xinchen Liu et al.ICML 2026
Builds on30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Inpaint-Anywhere: Zero-Shot Multi-Identity Inpainting with Efficient Diffusion TransformerJunsheng Luan, Lei Zhao, Wei XingAAAI 2026
- MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic ExpansionFei Peng, Junqiang Wu, Yan Li, Tingting Gao et al.ICCV 2025 · 1 citation
- MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout GuidanceXierui Wang, Siming Fu, Qihan Huang, Wanggui He et al.ICLR 2025
- EGGen: Image Generation with Multi-entity Prior Learning through Entity GuidanceZhenhong Sun, Junyan Wang, Zhiyu Tan, Daoyi Dong et al.ACM MM 2024 · 4 citations
- Insert Anything: Image Insertion via In-Context Editing in DiTWensong Song, Hong Jiang, Zongxing Yang, Zheqiao Cheng et al.AAAI 2026 · 1 citation
