JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment
Jiaxiang Shang, Yu Zeng, Xin Qiao, Xin Wang, Runze Zhang, Guangyuan Sun, Vishal Patel, Hongbo Fu
Abstract
Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues for cross-identity scenarios, i.e., when the source and the driving subjects are different. Current self-supervised face reconstruction methods also demonstrate impressive results. However, these methods do not handle large expressions well, since their training data lacks samples of large expressions, and 2D facial attributes are inaccurate on such samples. To mitigate the above problems, we propose to explore the inner connection between the two tasks, i.e., using face reconstruction to provide sufficient 3D information for reenactment, and synthesizing videos paired with captured face model parameters through face reenactment to enhance the expression module of face reconstruction. In particular, we propose a novel cascade framework named JR2Net for Joint Face Reconstruction and Reenactment, which begins with the training of a coarse reconstruction network, followed by a 3D-aware face reenactment network based on the coarse reconstruction results. In the end, we train an expression tracking network based on our synthesized videos composed by image-face model parameter pairs. Such an expression tracking network can further enhance the coarse face reconstruction. Extensive experiments show that our JR2Net outperforms the state-of-the-art methods on several face reconstruction and reenactment benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 710 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen TargetsSungjoo Ha, Martin Kersner, Beomsu Kim, Seokjun Seo et al.AAAI 2020 · 184 citations
- EMOCA: Emotion Driven Monocular Face Capture and AnimationRadek Danecek, Michael J. Black, Timo BolkartCVPR 2022 · 180 citations
Related papers
- Mesh Guided One-shot Face Reenactment Using Graph Convolutional NetworksGuangming Yao, Yi Yuan, Tianjia Shao, Kun ZhouACM MM 2020 · 42 citations
- Dual-Generator Face ReenactmentGee-Sern Hsu, Chun-Hung Tsai, Hung-Yi WuCVPR 2022 · 39 citations
- FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression FeaturesAndre Rochow, Max Schwarz, Sven BehnkeCVPR 2024 · 17 citations
- DeepFaceFlow: In-the-Wild Dense 3D Facial Motion EstimationMohammad Rami Koujan, Anastasios Roussos, Stefanos ZafeiriouCVPR 2020
- FReeNet: Multi-Identity Face ReenactmentJiangning Zhang, Xianfang Zeng, Mengmeng Wang, Yusu Pan et al.CVPR 2020
