LipFormer: High-fidelity and Generalizable Talking Face Generation with A Pre-learned Facial Codebook
Jiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang, Yujun Shen, Deli Zhao, Jingren Zhou
摘要
Generating a talking face video from the input audio sequence is a practical yet challenging task. Most existing methods either fail to capture fine facial details or need to train a specific model for each identity. We argue that a codebook pre-learned on high-quality face images can serve as a useful prior that facilitates high-fidelity and generalizable talking head synthesis. Thanks to the strong capability of the codebook in representing face textures, we simplify the talking face generation task as finding proper lip-codes to characterize the variation of lips during portrait talking. To this end, we propose LipFormer, a Transformer-based framework to model the audio-visual coherence and predict the lip-codes sequence based on input audio features. We further introduce an adaptive face warping module, which helps warp the reference face to the target pose in the feature space, to alleviate the difficulty of lip-code prediction under different poses. By this means, LipFormer can make better use of prelearned priors in images and is robust to posture change. Extensive experiments show that LipFormer can produce more realistic talking face videos compared to previous methods and faithfully generalize to unseen identities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression FeaturesAndre Rochow, Max Schwarz, Sven BehnkeCVPR 2024 · 被引用 17 次
- FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationShuai Tan, Bin Ji, Ye PanCVPR 2024 · 被引用 17 次
- FaceComposer: A Unified Model for Versatile Facial Content CreationJiayu Wang, Kang Zhao, Yifeng Ma, Shiwei Zhang 等NeurIPS 2023 · 被引用 14 次
- FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioChao Xu, Yang Liu, Jiazheng Xing, Weida Wang 等CVPR 2024 · 被引用 11 次
它引用的顶会 Paper17
- Implicit Geometric Regularization for Learning ShapesAmos Gropp, Lior Yariv, Niv Haim, Matan Atzmon 等ICML 2020 · 被引用 1,001 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- Towards Robust Blind Face Restoration with Codebook Lookup TransformerShangchen Zhou, Kelvin C. K. Chan, Chongyi Li, Chen Change LoyNeurIPS 2022 · 被引用 431 次
相关 Paper
- Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsWeizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei 等CVPR 2023
- GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific AdaptationWentao Hu, Shunkai Li, Ziqiao Peng, Haoxian Zhang 等ICCV 2025
- FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion SynthesisMengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan 等ACM MM 2025 · 被引用 9 次
- One-Shot Talking Face Generation from Single-Speaker Audio-Visual Correlation LearningSuzhen Wang, Lincheng Li, Yu Ding, Xin YuAAAI 2022 · 被引用 142 次
- GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsZiqi Zhou, Weize Quan, Hailin Shi, Wei Li 等AAAI 2025 · 被引用 1 次
