Optimizing ID Consistency in Multimodal Large Models: Facial Restoration via Alignment, Entanglement, and Disentanglement
Yuran Dong, Hang Dai, Mang Ye
Abstract
Multimodal editing large models have demonstrated powerful editing capabilities across diverse tasks. However, a persistent and long-standing limitation is the decline in facial identity (ID) consistency during realistic portrait editing. Due to the human eye’s high sensitivity to facial features, such inconsistency significantly hinders the practical deployment of these models. Current facial ID preservation methods struggle to achieve consistent restoration of both facial identity and edited element IP due to Cross-source Distribution Bias and Cross-source Feature Contamination. To address these issues, we propose EditedID, an Alignment-Disentanglement-Entanglement framework for robust identity-specific facial restoration. By systematically analyzing diffusion trajectories, sampler behaviors, and attention properties, we introduce three key components: 1) Adaptive mixing strategy that aligns cross-source latent representations throughout the diffusion process. 2) Hybrid solver that disentangles source-specific identity attributes and details. 3) Attentional gating mechanism that selectively entangles visual elements. Extensive experiments show that EditedID achieves state-of-the-art performance in preserving original facial ID and edited element IP consistency. As a training-free and plug-and-play solution, it establishes a new benchmark for practical and reliable single/multi-person facial identity restoration in open-world settings, paving the way for the deployment of multimodal editing large models in real-person editing scenarios. The code is available at https://github.com/NDYBSNDY/EditedID.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ef31ab4-25df-4b00-b25f-bee2059c0f4fBuilds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 670 citations
Related papers
- OmniPortrait: Fine-Grained Personalized Portrait Synthesis via Pivotal OptimizationDongxu Yue, Bo Lin, Yao Tang, Jiajun Liang et al.ICLR 2026
- VividFace: A Robost and High-Fidelity Video Face Swapping FrameworkHao Shao, Shulun Wang, Yang Zhou, Guanglu Song et al.NeurIPS 2025 · 4 citations
- UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image PersonalizationJunjie He, Yifeng Geng, Liefeng BoICCV 2025 · 2 citations
- Face2Diffusion for Fast and Editable Face PersonalizationKaede Shiohara, Toshihiko YamasakiCVPR 2024 · 13 citations
- Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation KnowledgeYiyang Cai, Zhengkai Jiang, Yulong Liu, Chunyang Jiang et al.NeurIPS 2025 · 2 citations
