Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
Tianqi Li, Ruobing Zheng, Minghui Yang, Jingdong Chen, Ming Yang
摘要
Recent advances in diffusion models have endowed talking head synthesis with subtle expressions and vivid head movements, but have also led to slow inference speed and insufficient control over generated results. To address these issues, we propose Ditto, a diffusion-based talking head framework that enables fine-grained controls and real-time inference. Specifically, we utilize an off-the-shelf motion extractor and devise a diffusion transformer to generate representations in a specific motion space. We optimize the model architecture and training strategy to address the issues in generating motion representations, including insufficient disentanglement between motion and identity, and large internal discrepancies within the representation. Besides, we employ diverse conditional signals while establishing a mapping between motion representation and facial semantics, enabling control over the generation process and correction of the results. Moreover, we jointly optimize the holistic framework to enable streaming processing, real-time inference, and low first-frame delay, offering functionalities crucial for interactive applications such as AI assistants. Extensive experimental results demonstrate that Ditto generates compelling talking head videos and exhibits superiority in both controllability and real-time performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingZhongjian Wang, Peng Zhang, Jinwei Qi, Yuan Wang 等NeurIPS 2025 · 被引用 12 次
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu 等CVPR 2026 · 被引用 8 次
- Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake DetectionTianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He 等CVPR 2026 · 被引用 5 次
- REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming DistillationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 等ICML 2026 · 被引用 4 次
- Versatile Multimodal Controls for Expressive Talking Human AnimationZheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
相关 Paper
- ExpPortrait: Expressive Portrait Generation via Personalized RepresentationJunyi Wang, Yudong Guo, Boyang Guo, Shengming Yang 等CVPR 2026
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head GenerationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 等AAAI 2026 · 被引用 1 次
- FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion ModelZiyu Yao, Xuxin Cheng, Zhiqi HuangACM MM 2024 · 被引用 5 次
- DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationShuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 等CVPR 2023
- MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head GenerationSeyeon Kim, Siyoon Jin, Jihye Park, Kihong Kim 等AAAI 2025 · 被引用 12 次
