Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
Fei Shen, Cong Wang, Junyao Gao, Qin Guo, Jisheng Dang, Jinhui Tang, Tat-Seng Chua
Abstract
Recent advances in conditional diffusion models have shown promise for generating realistic Talk-ingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the Motion-priors Conditional Diffusion Model (MCDM), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also release the TalkingFace-Wild dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term Talk-ingFace generation. Code, models, and datasets will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a197bfaf-2ec6-4a54-9763-1c31873a3a4eCited by top-tier papers12
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie et al.CVPR 2026 · 14 citations
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationCong Wang, Zexuan Deng, Zhiwei Jiang, Yafeng Yin et al.NeurIPS 2025 · 13 citations
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency RewardsShulei Wang, Longhui Wei, Xin He, Jianbo Ouyang et al.CVPR 2026 · 7 citations
- WildActor: Unconstrained Identity-Preserving Video GenerationQin Guo, Tianyu Yang, Xuanhua He, Fei Shen et al.ICML 2026 · 4 citations
- Ensembling Diffusion Models via Adaptive Feature AggregationCong Wang, Kuan Tian, Yonghang Guan, Fei Shen et al.ICLR 2025 · 3 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian et al.ACM MM 2024 · 3 citations
- MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head GenerationSeyeon Kim, Siyoon Jin, Jihye Park, Kihong Kim et al.AAAI 2025 · 12 citations
- Talking Head Generation with Probabilistic Audio-to-Visual Diffusion PriorsZhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang et al.ICCV 2023 · 65 citations
- GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsZiqi Zhou, Weize Quan, Hailin Shi, Wei Li et al.AAAI 2025 · 1 citation
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
