Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
Chanhyuk Choi, Taesoo Kim, Donggyu Lee, Siyeol Jung, Taehwan Kim
Abstract
Talking face generation has gained significant attention as a core application of generative models.To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role.However, existing approaches often limit expressive flexibility and struggle to generate extended emotions.Label-based methods represent emotions with discrete categories, which fail to capture a wide range of emotions.Audio-based methods can leverage emotionally rich speech signals—and even benefit from expressive text-to-speech (TTS) synthesis—but they fail to express the target emotions because emotions and linguistic contents are entangled in emotional speeches.Images-based methods, on the other hand, rely on target reference images to guide emotion transfer, yet they require high-quality frontal views and face challenges in acquiring reference data for extended emotions (e.g., sarcasm).To address these limitations, we propose Cross-Modal Emotion Transfer (C-MET), a novel approach that generates facial expressions based on speeches by modeling emotion semantic vectors between speech and visual feature spaces.C-MET leverages a large-scale pretrained audio encoder and a disentangled facial expression encoder to learn emotion semantic vectors that represent the difference between two different emotional embeddings across modalities.Extensive experiments on the MEAD and CREMA-D datasets demonstrate that our method improves emotion accuracy by 14% over state-of-the-art methods, while generating expressive talking face videos—even for unseen extended emotions. All source code and checkpoint will be released upon acceptance, including video samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsZhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li et al.AAAI 2025 · 197 citations
Related papers
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningChao Xu, Junwei Zhu, Jiangning Zhang, Yue Han et al.CVPR 2023
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait GenerationWeipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu et al.ACM MM 2025 · 4 citations
