FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio
Chao Xu, Yang Liu, Jiazheng Xing, Weida Wang, Mingze Sun, Jun Dan, Tianxin Huang, Siyuan Li, Zhi-Qi Cheng, Ying Tai, Baigui Sun
摘要
In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating vari-ous dynamically audio-consistent talking faces, termed Lis-tening and Imagining, into the task of high-fidelity diverse talking faces generation from a single audio. Specifically, it involves two critical challenges: one is to effectively de-couple identity, content, and emotion from entangled au-dio, and the other is to maintain intra-video diversity and inter- video consistency. To tackle the issues, we first dig out the intricate relationships among facial factors and sim-plify the decoupling process, tailoring a Progressive Audio Disentanglement for accurate facial geometry and seman-tics learning, where each stage incorporates a customized training module responsible for a specific factor. Secondly, to achieve visually diverse and audio-synchronized animation solely from input audio within a single model, we intro-duce the Controllable Coherent Frame generation, which involves the flexible integration of three trainable adapters with frozen Latent Diffusion Models (LDMs) to focus on maintaining facial geometry and semantics, as well as texsture and temporal coherence between frames. In this way, we inherit high-quality diverse generation from LDMs while significantly improving their controllability at a low training cost. Extensive experiments demonstrate the flexibility and effectiveness of our method in handling this paradigm. The codes will be released at FaceChain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong 等ICCV 2025 · 被引用 5 次
- POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position SearchChong-Yang Xiang, Jun-Yan He, Zhi-Qi Cheng, Xiao Wu 等AAAI 2025 · 被引用 3 次
- FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different StylesTian-Hao Zhang, Jiawei Zhang, Jun Wang, Xinyuan Qian 等AAAI 2025 · 被引用 2 次
- Learning to Decouple the Lights for 3D Face Texture ModelingTianxin Huang, Zhenyu Zhang, Ying Tai, Gim Hee LeeNeurIPS 2024 · 被引用 2 次
- Synchronized Video-to-Audio Generation via Mel Quantization-Continuum DecompositionJuncheng Wang, Chao Xu, Cheng Yu, Lei Shang 等CVPR 2025
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsZiqi Zhou, Weize Quan, Hailin Shi, Wei Li 等AAAI 2025 · 被引用 1 次
- Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait GenerationWeipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu 等ACM MM 2025 · 被引用 4 次
- ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian 等ACM MM 2024 · 被引用 3 次
- DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationShuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 等CVPR 2023
- FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion SynthesisMengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan 等ACM MM 2025 · 被引用 9 次
