XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
Bowen Chen, Brynn zhao, Haomiao Sun, Li Chen, Xu Wang, Daniel K. Du, Xinglong Wu
Abstract
Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of Diffusion Transformers (DiTs). Many approaches introduce artifacts or suffer from attribute entanglement. To overcome these challenges, we propose a novel multi-subject controlled generation model XVerse. By transforming reference images into offsets for token-specific text-stream modulation, XVerse allows for precise and independent control for specific subject without disrupting image latents or features. Consequently, XVerse offers high-fidelity, editable multi-subject image synthesis with robust control over individual subject characteristics and semantic attributes. This advancement significantly improves personalized and complex scene generation capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance GenerationRuihang Xu, Dewei Zhou, Fan Ma, Yi YangICLR 2026 · 19 citations
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and DisentanglementDong She, Siming Fu, Mushui Liu, Qiaoqiao Jin et al.ICLR 2026 · 13 citations
- WithAnyone: Toward Controllable and ID Consistent Image GenerationHengyuan Xu, Wei Cheng, Peng Xing, Yixiao Fang et al.ICLR 2026 · 12 citations
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image CompositionXinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo et al.CVPR 2026 · 10 citations
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation ModelingYuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu et al.CVPR 2026 · 9 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- TokenVerse: Versatile Multi-concept Personalization in Token Modulation SpaceDaniel Garibi, Shahar Yadin, Roni Paiss, Omer Tov et al.SIGGRAPH 2025 · 12 citations
- MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic ExpansionFei Peng, Junqiang Wu, Yan Li, Tingting Gao et al.ICCV 2025 · 1 citation
- Less-to-More Generalization: Unlocking More Controllability by In-Context GenerationShaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng et al.ICCV 2025 · 11 citations
- AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and PoseHuichao Zhang, Bowen Chen, Hao Yang, Liao Qu et al.AAAI 2024 · 73 citations
- DynASyn: Multi-Subject Personalization Enabling Dynamic Action SynthesisYongjin Choi, Chanhun Park, Seung Jun BaekAAAI 2025 · 3 citations
