MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation
Yinuo Wang, Yanbo Fan, Xuan Wang, Boyao Zhou, Yu Guo, Yujun Shen, Fei Wang
摘要
Dyadic interactive head generation aims to synthesize realistic head motions that respond both verbally and nonverbally to an interlocutor in real-time conversation. The existing works often focus on offline scenarios, and struggle with a shallow understanding of the multimodal conversational context while also lacking long-term coherence. To address these limitations, we propose MimicTalker, a novel method for producing real-time, contextually-aware, and long-term consistent interactive head motions. To this end, we propose a Multimodal Interactive Context Extraction (MICE) module to capture both instantaneous and longterm multimodal interactive information from the interlocutor. To enhance in-depth conversational understanding, we propose a Semantic-enhanced Dynamic Interaction (SDI) module to integrate the intentions and topics of the conversation, which are automatically extracted through an LLMbased analyzer. Further, we propose a semantic-guided Motion Style Memory (MSM) mechanism, enabling the longterm motion consistency throughout the conversation. We conduct experiments on both short conversational segments (25 seconds) and extended dialogues (6 minutes), and the comprehensive experiments demonstrate that our method significantly outperforms existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 等CVPR 2022 · 被引用 218 次
相关 Paper
- MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head GenerationSeyeon Kim, Siyoon Jin, Jihye Park, Kihong Kim 等AAAI 2025 · 被引用 12 次
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen 等AAAI 2026 · 被引用 2 次
- MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationHyung Kyu Kim, Sangmin Lee, Hak Gu KimICCV 2025 · 被引用 1 次
- Talking Together: Synthesizing Co-Located 3D Conversations from AudioMengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu 等CVPR 2026
- Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion GenerationBohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao 等ACM MM 2024 · 被引用 26 次
