UniFace: A fied ine-grained Understanding and Generation Model
Junzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu, Linrui Xu, TingTing Long, Chun Fan, Ming Li, Hehe Fan, Jun Liu, Shuicheng YAN
Abstract
Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: (1) fragmentation development, with existing methods failing to unify understanding and generation into a single one, hindering the way to artificial general intelligence. (2) lack of fine-grained facial attributes, which are crucial for high-fidelity applications. To handle those issues, we propose UniFace, the first UMM specifically tailored for fine-grained face understanding and generation. First, we introduce a novel theoretical framework with a Dual Discrete Diffusion (D3Diff) loss, unifying masked generative models with discrete score matching diffusion and leading to a more precise approximation of the negative log-likelihood. Moreover, this D3Diff significantly enhances the model's ability to synthesize high-fidelity facial details aligned with text input. Second, we propose a multi-level grouped Mixture-of-Experts architecture, adaptively incorporating the semantic and identity facial embeddings to complement the attribute forgotten phenomenon in representation evolvement. Finally, to this end, we construct UniFaceD-1M, a large-scale dataset comprising 130K fine-grained image-caption pairs and 1M visual question-answering pairs, spanning a much wider range of facial attributes than existing datasets. Extensive experiments demonstrate that UniFace outperforms existing models with a similar scale in both understanding and generation tasks, with 7.1% higher Desc-GPT and 6.6% higher VQA-score, respectively. Code is available in the supplementary materials.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73d91ac4-3694-40be-92c3-44830a71b44eBuilds on50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial GenerationYushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou et al.AAAI 2026
- MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationXiangdong Li, Ye Lou, Ao Gao, Wei Zhang et al.AAAI 2026
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete DiffusionLijiang Li, zuwei long, Yunhang Shen, Heting Gao et al.ICML 2026 · 7 citations
- A Rich Knowledge Space for Scalable Deepfake DetectionInho Jung, Hyeongjun Choi, Binh Minh Le, Hohyun Na et al.ICLR 2026
- FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and GenerationPeng Zhang, Wanggui He, Mushui Liu, Wenyi Xiao et al.AAAI 2026
