Audio-Driven Emotional Video Portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, Feng Xu
Abstract
Target ID-1 Sad Emotion Interpolation (a) (b) Happy Sad Figure 1: Audio-Driven Emotional Video Portraits. Given an audio clip and a target video, our Emotional Video Portraits (EVP) approach is capable of generating emotion-controllable talking portraits and change the emotion of them smoothly by interpolating at the latent space. (a) Generated video portraits with the same speech content but different emotions (i.e., contempt and sad). (b) Linear interpolation of the learned latent representation of emotions from sad to happy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59a3e154-c023-49bb-b4af-31f5eb20328dCited by top-tier papers46
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang et al.CVPR 2022 · 218 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation LearningSuzhen Wang, Lincheng Li, Yu Ding, Xin YuAAAI 2022 · 142 citations
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
Builds on2
Related papers
- FLOAT: Generative Motion Latent Flow Matching for Audio-Driven Talking PortraitTaekyung Ki, Dongchan Min, Gyeongsu ChaeICCV 2025 · 6 citations
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu et al.SIGGRAPH 2022 · 150 citations
- Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided DiffusionXingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang et al.ICML 2025
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking HeadsHaoyu Wang, Xiaozhe Xin, Xiaoyu Qin, Meiguang Jin et al.AAAI 2026 · 1 citation
