Learning Dense Correspondence for NeRF-Based Face Reenactment
Songlin Yang, Wei Wang, Yushi Lan, Xiangyu Fan, Bo Peng, Lei Yang, Jing Dong
Abstract
Face reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment in photo-realism and 3D consistency. However, establishing dense correspondence between different face NeRFs is non-trivial, because implicit representations lack ground-truth correspondence annotations like mesh-based 3D parametric models (e.g., 3DMM) with index-aligned vertexes. Although aligning 3DMM space with NeRF-based face representations can realize motion control, it is sub-optimal for their limited face-only modeling and low identity fidelity. Therefore, we are inspired to ask: Can we learn the dense correspondence between different NeRF-based face representations without a 3D parametric model prior? To address this challenge, we propose a novel framework, which adopts tri-planes as fundamental NeRF representation and decomposes face tri-planes into three components: canonical tri-planes, identity deformations, and motion. In terms of motion control, our key contribution is proposing a Plane Dictionary (PlaneDict) module, which efficiently maps the motion conditions to a linear weighted addition of learnable orthogonal plane bases. To the best of our knowledge, our framework is the first method that achieves one-shot multi-view face reenactment without a 3D parametric model prior. Extensive experiments demonstrate that we produce better results in fine-grained motion control and identity preservation than previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecc1dad6-5cb5-46a3-993c-d8d0e754c935Cited by top-tier papers3
- Generalizable and Animatable Gaussian Head AvatarXuangeng Chu, Tatsuya HaradaNeurIPS 2024 · 115 citations
- OMG-Avatar: One-shot Multi-LOD Gaussian Head AvatarJianqiang Ren, Lin Liu, Steven HoiCVPR 2026 · 1 citation
- Textured 3D Regenerative Morphing with 3D Diffusion PriorSonglin Yang, Yushi Lan, Honghua Chen, Xingang PanICCV 2025 · 1 citation
Builds on21
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image SynthesisJiatao Gu, Lingjie Liu, Peng Wang, Christian TheobaltICLR 2022 · 622 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
Related papers
- NOFA: NeRF-based One-shot Facial Avatar ReconstructionWangbo Yu, Yanbo Fan, Yong Zhang, Xuan Wang et al.SIGGRAPH 2023 · 38 citations
- VOODOO 3D: Volumetric Portrait Disentanglement for One-Shot 3D Head ReenactmentPhong Tran, Egor Zakharov, Long-Nhat Ho, Anh Tuan Tran et al.CVPR 2024 · 15 citations
- Parametric Implicit Face Representation for Audio-Driven Facial ReenactmentRicong Huang, Peiwen Lai, Yipeng Qin, Guanbin LiCVPR 2023
- Real-Time 3D-Aware Portrait Video RelightingZiqi Cai, Kaiwen Jiang, Shu-Yu Chen, Yu-Kun Lai et al.CVPR 2024
- MonoHuman: Animatable Human Neural Field from Monocular VideoZhengming Yu, Wei Cheng, Xian Liu, Wayne Wu et al.CVPR 2023
