Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation
Hui Fu, Zeqing Wang, Ke Gong, Keze Wang, Tianshui Chen, Haojie Li, Haifeng Zeng, Wenxiong Kang
Abstract
Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while neglecting to model the subject-specific speaking style, often resulting in unrealistic facial animations. To the best of our knowledge, this work makes the first attempt to explore the coupled information between the speaking style and the semantic content in facial motions. Specifically, we introduce an innovative speaking style disentanglement method, which enables arbitrary-subject speaking style encoding and leads to a more realistic synthesis of speech-driven facial animations. Subsequently, we propose a novel framework called Mimic to learn disentangled representations of the speaking style and content from facial motions by building two latent spaces for style and content, respectively. Moreover, to facilitate disentangled representation learning, we introduce four well-designed constraints: an auxiliary style classifier, an auxiliary inverse classifier, a content contrastive loss, and a pair of latent cycle losses, which can effectively contribute to the construction of the identity-related style space and semantic-related content space. Extensive qualitative and quantitative experiments conducted on three publicly available datasets demonstrate that our approach outperforms state-of-the-art methods and is capable of capturing diverse speaking styles for speech-driven 3D facial animation. The source code and supplementary video are publicly available at: https://zeqing-wang.github.io/Mimic/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e28b62a-1866-40b9-b0fb-37c68ef8477fCited by top-tier papers4
- PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesTianshun Han, Benjia Zhou, Ajian Liu, Yanyan Liang et al.ACM MM 2025 · 1 citation
- MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationHyung Kyu Kim, Sangmin Lee, Hak Gu KimICCV 2025 · 1 citation
- PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentBin Wang, Yang Xu, Huan Zhao, Hao Zhang et al.ACM MM 2025
- Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationHao Li, Ju Dai, Xin Zhao, Feng Zhou et al.CVPR 2025
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech GenerationDongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju HwangICML 2021 · 218 citations
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang et al.CVPR 2022 · 218 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
Related papers
- Imitator: Personalized Speech-driven 3D Facial AnimationBalamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker et al.ICCV 2023 · 98 citations
- Model See Model Do: Speech-Driven Facial Animation with Style ControlYifang Pan, Karan Singh, Luiz Gustavo HafemannSIGGRAPH 2025 · 2 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- EcoFace: Audio-Visual Emotional Co-Disentanglement Speech-Driven 3D Talking Face GenerationJiajian Xie, Shengyu Zhang, Mengze Li, Chengfei Lv et al.ICLR 2025
- CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorJinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun et al.CVPR 2023
