Expressive Talking Avatars
Ye Pan, Shuai Tan, Shengran Cheng, Qunfen Lin, Zijiao Zeng, Kenny Mitchell
Abstract
Stylized avatars are common virtual representations used in VR to support interaction and communication between remote collaborators. However, explicit expressions are notoriously difficult to create, mainly because most current methods rely on geometric markers and features modeled for human faces, not stylized avatar faces. To cope with the challenge of emotional and expressive generating talking avatars, we build the Emotional Talking Avatar Dataset which is a talking-face video corpus featuring 6 different stylized characters talking with 7 different emotions. Together with the dataset, we also release an emotional talking avatar generation method which enables the manipulation of emotion. We validated the effectiveness of our dataset and our method in generating audio based puppetry examples, including comparisons to state-of-the-art techniques and a user study. Finally, various applications of this method are discussed in the context of animating avatars in VR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Say Anything with Any StyleShuai Tan, Bin Ji, Yu Ding, Ye PanAAAI 2024 · 30 citations
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 17 citations
- Through the Eyes of Emotion: A Multi-faceted Eye Tracking Dataset for Emotion Recognition in Virtual RealityTongyun Yang, Bishwas Regmi, Lingyu Du, Andreas Bulling et al.UbiComp 2025 · 3 citations
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesShuai Tan, Bill Gong, Bin Ji, Ye PanICCV 2025 · 3 citations
- When Words Smile: Generating Diverse Emotional Facial Expressions from TextHaidong Xu, Meishan Zhang, Hao Ju, Zhedong Zheng et al.EMNLP 2025 · 2 citations
Builds on10
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelXinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu et al.SIGGRAPH 2022 · 150 citations
- EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face GenerationShuai Tan, Bin Ji, Ye PanICCV 2023 · 63 citations
- SPACE: Speech-driven Portrait Animation with Controllable ExpressionSiddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle et al.ICCV 2023 · 58 citations
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu et al.ICLR 2023 · 33 citations
Related papers
- MoEE: Mixture of Emotion Experts for Audio-Driven Portrait AnimationHuaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun et al.CVPR 2025
- Emotional Voice PuppetryYe Pan, Ruisi Zhang, Shengran Cheng, Shuai Tan et al.IEEE VR 2023 · 21 citations
- VASA-Rig: Audio-Driven 3D Facial Animation with 'Live' Mood Dynamics in Virtual RealityYe Pan, Chang Liu, Sicheng Xu, Shuai Tan et al.IEEE VR 2025 · 5 citations
- EmoFace: Audio-driven Emotional 3D Face AnimationChang Liu, Qunfen Lin, Zijiao Zeng, Ye PanIEEE VR 2024 · 16 citations
- Expressive Talking Head Generation with Granular Audio-Visual ControlBorong Liang, Yan Pan, Zhizhi Guo, Hang Zhou et al.CVPR 2022 · 114 citations
