Parametric Implicit Face Representation for Audio-Driven Facial Reenactment
Ricong Huang, Peiwen Lai, Yipeng Qin, Guanbin Li
Abstract
Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial landmarks or 3D face models) or implicit ones (e.g., Neural Radiance Fields), thus suffering from the trade-offs between interpretability and expressive power, hence between controllability and quality of the results. In this work, we break these trade-offs with our novel parametric implicit face representation and propose a novel audio-driven facial reenactment framework that is both controllable and can generate high-quality talking heads. Specifically, our parametric implicit representation parameterizes the implicit representation with interpretable parameters of 3D face models, thereby taking the best of both explicit and implicit methods. In addition, we propose several new techniques to improve the three components of our framework, including i) incorporating contextual information into the audio-to-expression parameters encoding; ii) using conditional image synthesis to parameterize the implicit representation and implementing it with an innovative tri-plane structure for efficient learning; iii) formulating facial reenactment as a conditional image inpainting problem and proposing a novel data augmentation technique to improve model generalizability. Extensive experiments demonstrate that our method can generate more realistic results than previous methods with greater fidelity to the identities and talking styles of speakers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4753d91-599d-4892-b91f-1b34aa6640adCited by top-tier papers4
- ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion ModelingQuanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu et al.NeurIPS 2024 · 21 citations
- Hierarchically Controlled Deformable 3D Gaussians for Talking Head SynthesisZhenhua Wu, Linxuan Jiang, Xiang Li, Chaowei Fang et al.AAAI 2025 · 2 citations
- SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking HeadFatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Josef Kittler et al.CVPR 2026
- LLM-driven Multimodal and Multi-Identity Listening Head GenerationPeiwen Lai, Weizhi Zhong, Yipeng Qin, Xiaohang Ren et al.CVPR 2025
Builds on12
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
Related papers
- Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar ReconstructionGuy Gafni, Justus Thies, Michael Zollhöfer, Matthias NießnerCVPR 2021
- OTAvatar: One-Shot Talking Face Avatar with Controllable Tri-Plane RenderingZhiyuan Ma, Xiangyu Zhu, Guojun Qi, Zhen Lei et al.CVPR 2023
- One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldWeichuang Li, Longhao Zhang, Dong Wang, Bin Zhao et al.CVPR 2023
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen et al.ACM MM 2024 · 28 citations
- Context-Aware Talking-Head Video EditingSonglin Yang, Wei Wang, Jun Ling, Bo Peng et al.ACM MM 2023 · 10 citations
