Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching Paradigm
Yufeng Cheng, wenxu wu, Shaojin Wu, Mengqi Huang, Fei Ding, Qian HE
摘要
Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-reference images, limiting the identity scalability of customization models. To address this, we present UMO, a Unified Multi-identity Optimization framework, designed to maintain high-fidelity identity preservation and alleviate identity confusion with scalability. With multi-to-multi matching paradigm, UMO reformulates multi-identity generation as a global assignment optimization problem and unleashes multi-identity consistency for existing image customization methods generally. To facilitate the training of UMO, we develop a customization dataset with multi-reference images, consisting of both synthesised and real parts. Additionally, we propose a new metric to measure identity confusion. Extensive experiments demonstrate that UMO not only improves identity consistency significantly, but also reduces identity confusion on several image customization methods, setting a new state-of-the-art among open-source methods along the dimension of identity preserving.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- OmniPortrait: Fine-Grained Personalized Portrait Synthesis via Pivotal OptimizationDongxu Yue, Bo Lin, Yao Tang, Jiajun Liang 等ICLR 2026
- SIGMA-Gen: Structure and Identity Guided Multi-Subject Assembly for Image GenerationOindrila Saha, Vojtech Krs, Radomir Mech, Subhransu Maji 等ICLR 2026 · 被引用 5 次
- UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image GenerationDanning Zhang, Yijing Lin, Shuhan Zhuang, Mengqi Huang 等ICML 2026
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video GenerationRun Ling, Ke Cao, Jian Lu, Ao Ma 等AAAI 2026 · 被引用 4 次
- Unified Customized Generation by Disentangled Reward ModelingShaojin Wu, Mengqi Huang, Yufeng Cheng, Wenxu Wu 等CVPR 2026
