EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face Generation
Shuai Tan, Bin Ji, Ye Pan
Abstract
Synthesizing expression is essential to create realistic talking faces. Previous works consider expressions and mouth shapes as a whole and predict them solely from audio inputs. However, the limited information contained in audio, such as phonemes and coarse emotion embedding, may not be suitable as the source of elaborate expressions. Besides, since expressions are tightly coupled to lip motions, generating expression from other sources is tricky and always neglects expression performed on mouth region, leading to inconsistency between them. To tackle the issues, this paper proposes Emotional Motion Memory Net (EMMN) that synthesizes expression overall on the talking face via emotion embedding and lip motion instead of the sole audio. Specifically, we extract emotion embedding from audio and design Motion Reconstruction module to decompose ground truth videos into mouth features and expression features before training, where the latter encode all facial factors about expression. During training, the emotion embedding and mouth features are used as keys, and the corresponding expression features are used as values to create key-value pairs stored in the proposed Motion Memory Net. Hence, once the audio-relevant mouth features and emotion embedding are individually predicted from audio at inference time, we treat them as a query to retrieve the best-matching expression features, performing expression overall on the face and thus avoiding inconsistent results. Extensive experiments demonstrate that our method can generate high-quality talking face videos with accurate lip movements and vivid expressions on unseen subjects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69563029-cd47-4337-b8ff-45c787c84ddfCited by top-tier papers23
- Say Anything with Any StyleShuai Tan, Bin Ji, Yu Ding, Ye PanAAAI 2024 · 30 citations
- MimicTalk: Mimicking a personalized and expressive 3D talking face in minutesZhenhui Ye, Tianyun Zhong, Yi Ren, Ziyue Jiang et al.NeurIPS 2024 · 28 citations
- Expressive Talking AvatarsYe Pan, Shuai Tan, Shengran Cheng, Qunfen Lin et al.IEEE VR 2024 · 20 citations
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 17 citations
- MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion TransformerShurong Yang, Huadong Li, Juhao Wu, Minhao Jing et al.AAAI 2025 · 12 citations
Builds on12
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li et al.ICCV 2021 · 284 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
Related papers
- SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip MemorySe Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi et al.AAAI 2022 · 110 citations
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu et al.SIGGRAPH 2022 · 150 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationHyung Kyu Kim, Sangmin Lee, Hak Gu KimICCV 2025 · 1 citation
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang et al.CVPR 2023
