LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces From Video Using Pose and Lighting Normalization
Avisek Lahiri, Vivek Kwatra, Christian Früh, John Lewis, Chris Bregler
摘要
In this paper, we present a video-based learning framework for animating personalized 3D talking faces from audio. We introduce two training-time data normalizations that significantly improve data sample efficiency. First, we isolate and represent faces in a normalized space that decouples 3D geometry, head pose, and texture. This decomposes the prediction problem into regressions over the 3D face shape and the corresponding 2D texture atlas. Second, we leverage facial symmetry and approximate albedo constancy of skin to isolate and remove spatio-temporal lighting variations. Together, these normalizations allow simple networks to generate high fidelity lip-sync videos under novel ambient illumination while training with just a single speaker-specific video. Further, to stabilize temporal dynamics, we introduce an auto-regressive approach that conditions the model on its previous visual state. Human ratings and objective metrics demonstrate that our method outperforms contemporary state-of-the-art audio-driven video reenactment benchmarks in terms of realism, lip-sync and visual quality scores. We illustrate several applications enabled by our framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 等CVPR 2022 · 被引用 218 次
- One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation LearningSuzhen Wang, Lincheng Li, Yu Ding, Xin YuAAAI 2022 · 被引用 142 次
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan 等AAAI 2023 · 被引用 135 次
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan 等AAAI 2023 · 被引用 106 次
- Imitator: Personalized Speech-driven 3D Facial AnimationBalamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker 等ICCV 2023 · 被引用 98 次
它引用的顶会 Paper1
相关 Paper
- PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentBin Wang, Yang Xu, Huan Zhao, Hao Zhang 等ACM MM 2025
- Context-Aware Talking-Head Video EditingSonglin Yang, Wei Wang, Jun Ling, Bo Peng 等ACM MM 2023 · 被引用 10 次
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei 等CVPR 2023
- Talking Head Generation with Probabilistic Audio-to-Visual Diffusion PriorsZhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang 等ICCV 2023 · 被引用 65 次
- GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific AdaptationWentao Hu, Shunkai Li, Ziqiao Peng, Haoxian Zhang 等ICCV 2025
