UniFace: A fied ine-grained Understanding and Generation Model
Junzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu, Linrui Xu, TingTing Long, Chun Fan, Ming Li, Hehe Fan, Jun Liu, Shuicheng YAN
摘要
Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: (1) fragmentation development, with existing methods failing to unify understanding and generation into a single one, hindering the way to artificial general intelligence. (2) lack of fine-grained facial attributes, which are crucial for high-fidelity applications. To handle those issues, we propose UniFace, the first UMM specifically tailored for fine-grained face understanding and generation. First, we introduce a novel theoretical framework with a Dual Discrete Diffusion (D3Diff) loss, unifying masked generative models with discrete score matching diffusion and leading to a more precise approximation of the negative log-likelihood. Moreover, this D3Diff significantly enhances the model's ability to synthesize high-fidelity facial details aligned with text input. Second, we propose a multi-level grouped Mixture-of-Experts architecture, adaptively incorporating the semantic and identity facial embeddings to complement the attribute forgotten phenomenon in representation evolvement. Finally, to this end, we construct UniFaceD-1M, a large-scale dataset comprising 130K fine-grained image-caption pairs and 1M visual question-answering pairs, spanning a much wider range of facial attributes than existing datasets. Extensive experiments demonstrate that UniFace outperforms existing models with a similar scale in both understanding and generation tasks, with 7.1% higher Desc-GPT and 6.6% higher VQA-score, respectively. Code is available in the supplementary materials.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial GenerationYushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou 等AAAI 2026
- MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label GenerationXiangdong Li, Ye Lou, Ao Gao, Wei Zhang 等AAAI 2026
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete DiffusionLijiang Li, zuwei long, Yunhang Shen, Heting Gao 等ICML 2026 · 被引用 7 次
- A Rich Knowledge Space for Scalable Deepfake DetectionInho Jung, Hyeongjun Choi, Binh Minh Le, Hohyun Na 等ICLR 2026
- FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and GenerationPeng Zhang, Wanggui He, Mushui Liu, Wenyi Xiao 等AAAI 2026
