FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
Mengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan, Yunpeng Zhang, Yonggang Qi, Kun Zhao, Mu Xu
摘要
Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address these limitations, we propose a novel framework that leverages a pretrained video diffusion Transformer model to generate high-fidelity, coherent talking portraits with controllable motion dynamics. At the core of our work is a dual-stage audio-visual alignment strategy. In the first stage, we employ a clip-level training scheme to establish coherent global motion by aligning audio-driven dynamics across the entire scene, including the reference portrait, contextual objects, and background. In the second stage, we refine lip movements at the frame level using a lip-tracing mask, ensuring precise synchronization with audio signals. To preserve identity without compromising motion flexibility, we replace the commonly used reference network with a lightweight cross-attention module that effectively maintains facial consistency throughout the video. Furthermore, we integrate a motion intensity modulation module that explicitly controls facial keypoints and body joint trajectories, enabling fine-grained manipulation of portrait movements beyond mere lip motion. Extensive experimental results show that our proposed approach achieves higher quality with better realism, coherence, motion intensity, and identity preservation. Our demo, code, models can be found on this page: https://fantasy-amap.github.io/fantasy-talking/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Let Them Talk: Audio-Driven Multi-Person Conversational Video GenerationZhe Kong, Feng Gao, Yong Zhang, Zhuoliang Kang 等NeurIPS 2025 · 被引用 73 次
- SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human GenerationYouliang Zhang, Zhaoyang Li, Duomin Wang, jiahe zhang 等ICLR 2026 · 被引用 30 次
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang 等ICLR 2026 · 被引用 26 次
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen 等CVPR 2026 · 被引用 26 次
- EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human AnimationRang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng 等AAAI 2026 · 被引用 24 次
它引用的顶会 Paper23
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- RealPortrait: Realistic Portrait Animation with Diffusion TransformersZejun Yang, Huawei Wei, Zhisheng WangAAAI 2025 · 被引用 2 次
- Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion TransformerJiahao Cui, Hui Li, Yun Zhan, Hanlin Shang 等CVPR 2025
- SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking HeadFatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Josef Kittler 等CVPR 2026
- FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion ModelZiyu Yao, Xuxin Cheng, Zhiqi HuangACM MM 2024 · 被引用 5 次
- GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsZiqi Zhou, Weize Quan, Hailin Shi, Wei Li 等AAAI 2025 · 被引用 1 次
