3D Face Tracking from 2D Video through Iterative Dense UV to Image Flow
Felix Taubner, Prashant Raina, Mathieu Tuli, Eu Wern Teh, Chul Lee, Jinmiao Huang
Abstract
When working with 3D facial data, improving fidelity and avoiding the uncanny valley effect is critically dependent on accurate 3D facial performance capture. Because such methods are expensive and due to the widespread availability of 2D videos, recent methods have focused on how to perform monocular 3D face tracking. However, these methods often fall short in capturing precise facial movements due to limitations in their network architecture, training, and evaluation processes. Addressing these challenges, we propose a novel face tracker, FlowFace, that in-troduces an innovative 2D alignment network for dense pervertex alignment. Unlike prior work, FlowFace is trained on high-quality 3D scan annotations rather than weak supervision or synthetic data. Our 3D model fitting module Jointly fits a 3D face model from one or many observations, integrating existing neutral shape priors for enhanced identity and expression disentanglement and per-vertex de-formations for detailed facial feature reconstruction. Additionally, we propose a novel metric and benchmark for assessing tracking accuracy. Our method exhibits superior performance on both custom and publicly available bench-marks. We further validate the effectiveness of our tracker by generating high-quality 3D data from 2D videos, which leads to performance gains on downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c4029df-263e-4a5e-9b0c-bae0958fb559Cited by top-tier papers9
- Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face ReconstructionSimon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito et al.ICLR 2026 · 24 citations
- Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic CorrespondencePanagiotis Paraskevas Filntisis, George Retsinas, Radek Danecek, Vanessa Sklyarova et al.CVPR 2026 · 3 citations
- SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D GaussiansLiam Schoneveld, Zhe Chen, Davide Davoli, Jiapeng Tang et al.ICCV 2025 · 3 citations
- PhysHead: Simulation-Ready Gaussian Head AvatarsBerna Kabadayi, Vanessa Sklyarova, Wojciech Zielonka, Justus Thies et al.CVPR 2026 · 1 citation
- Feed-forward Gaussian Registration for Head Avatar Creation and EditingMalte Prinzler, Paulo F. U. Gotardo, Siyu Tang, Timo BolkartCVPR 2026 · 1 citation
Builds on16
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang et al.CVPR 2022 · 218 citations
- EMOCA: Emotion Driven Monocular Face Capture and AnimationRadek Danecek, Michael J. Black, Timo BolkartCVPR 2022 · 180 citations
- Neural Head Avatars from Monocular RGB VideosPhilip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother et al.CVPR 2022 · 173 citations
Related papers
- Accurate 3D Face Reconstruction with Facial Component TokensTianke Zhang, Xuangeng Chu, Yunfei Liu, Lijian Lin et al.ICCV 2023 · 38 citations
- DeepFaceFlow: In-the-Wild Dense 3D Facial Motion EstimationMohammad Rami Koujan, Anastasios Roussos, Stefanos ZafeiriouCVPR 2020
- Face Video Deblurring Using 3D Facial PriorsWenqi Ren, Jiaolong Yang, Senyou Deng, David P. Wipf et al.ICCV 2019 · 52 citations
- DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsXiaoxi Liang, Yanbo Fan, Qiya Yang, Xuan Wang et al.ICCV 2025 · 2 citations
- JR2Net: Joint Monocular 3D Face Reconstruction and ReenactmentJiaxiang Shang, Yu Zeng, Xin Qiao, Xin Wang et al.AAAI 2023 · 4 citations
