SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking Faces
Ziqiao Peng, Yihao Luo, Yue Shi, Hao Xu, Xiangyu Zhu, Hongyan Liu, Jun He, Zhaoxin Fan
摘要
Speech-driven 3D face animation technique, extending its applications to various multimedia fields.Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression models solely driven by data face several essential problems, such as difficulties in accessing precise labels and domain gaps between different modalities, leading to unsatisfactory results lacking precision and coherence.To enhance the visual accuracy of generated lip movement while reducing the dependence on labeled data, we propose a novel framework SelfTalk, by involving self-supervision in a cross-modals network system to learn 3D talking faces. The framework constructs a network system consisting of three modules: facial animator, speech recognizer, and lip-reading interpreter. The core of SelfTalk is a commutative training diagram that facilitates compatible features exchange among audio, text, and lip shape, enabling our models to learn the intricate connection between these factors. The proposed framework leverages the knowledge learned from the lip-reading interpreter to generate more plausible lip shapes. Extensive experiments and user studies demonstrate that our proposed approach achieves state-of-the-art performance both qualitatively and quantitatively. We recommend watching the supplementary video 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 等CVPR 2024 · 被引用 65 次
- Media2Face: Co-speech Facial Animation Generation With Multi-Modality GuidanceQingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin 等SIGGRAPH 2024 · 被引用 40 次
- OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersZiqiao Peng, Jiwen Liu, Haoxian Zhang, Xiaoqiang Liu 等NeurIPS 2025 · 被引用 30 次
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen 等ACM MM 2024 · 被引用 28 次
- MMHead: Towards Fine-grained Multi-modal 3D Facial AnimationSijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan 等ACM MM 2024 · 被引用 17 次
它引用的顶会 Paper19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre 等ICCV 2021 · 被引用 272 次
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 等CVPR 2022 · 被引用 218 次
相关 Paper
- CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorJinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun 等CVPR 2023
- Semi-supervised Speech-driven 3D Facial Animation via Cross-modal EncodingPeiji Yang, Huawei Wei, Yicheng Zhong, Zhisheng WangICCV 2023 · 被引用 1 次
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja 等ACM MM 2025 · 被引用 3 次
- SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningXiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang 等ACM MM 2024 · 被引用 4 次
- Speech Driven Tongue AnimationSalvador Medina, Denis Tomè, Carsten Stoll, Mark Tiede 等CVPR 2022 · 被引用 14 次
