Accurate 3D Face Reconstruction with Facial Component Tokens
Tianke Zhang, Xuangeng Chu, Yunfei Liu, Lijian Lin, Zhendong Yang, Zhengzhuo Xu, Chengkun Cao, Fei Yu, Changyin Zhou, Chun Yuan, Yu Li
Abstract
Accurately reconstructing 3D faces from monocular images and videos is crucial for various applications, such as digital avatar creation. However, the current deep learning-based methods face significant challenges in achieving accurate reconstruction with disentangled facial parameters and ensuring temporal stability in single-frame methods for 3D face tracking on video data. In this paper, we propose TokenFace, a transformer-based monocular 3D face reconstruction model. TokenFace uses separate tokens for different facial components to capture information about different facial parameters and employs temporal transformers to capture temporal information from video data. This design can naturally disentangle different facial components and is flexible to both 2D and 3D training data. Trained on hybrid 2D and 3D data, our model shows its power in accurately reconstructing faces from images and producing stable results for video data. Experimental results on popular benchmarks NoWand Stirling demonstrate that TokenFace achieves state-of-the-art performance, outperforming existing methods on all metrics by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- 3D Facial Expressions through Analysis-by-Neural-SynthesisGeorge Retsinas, Panagiotis Paraskevas Filntisis, Radek Danecek, Victoria Fernández Abrevaya et al.CVPR 2024 · 26 citations
- Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face ReconstructionSimon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito et al.ICLR 2026 · 24 citations
- Learning a Generalized Physical Face Model From DataLingchen Yang, Gaspard Zoss, Prashanth Chandran, Markus Gross et al.SIGGRAPH 2024 · 10 citations
- GUAVA: Generalizable Upper Body 3D Gaussian AvatarDongbin Zhang, Yunfei Liu, Lijian Lin, Ye Zhu et al.ICCV 2025 · 9 citations
- Makeup Prior Models for 3D Facial Makeup Estimation and ApplicationsXingchao Yang, Takafumi Taketomi, Yuki Endo, Yoshihiro KanamoriCVPR 2024 · 7 citations
Related papers
- AVFace: Towards Detailed Audio-Visual 4D Face ReconstructionAggelina Chatziagapi, Dimitris SamarasCVPR 2023
- 3D Face Tracking from 2D Video through Iterative Dense UV to Image FlowFelix Taubner, Prashant Raina, Mathieu Tuli, Eu Wern Teh et al.CVPR 2024 · 6 citations
- FlashAvatar: High-Fidelity Head Avatar with Efficient Gaussian EmbeddingJun Xiang, Xuan Gao, Yudong Guo, Juyong ZhangCVPR 2024 · 51 citations
- Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent TokensSen Yang, Wen Heng, Gang Liu, Guozhong Luo et al.ICLR 2023 · 4 citations
- TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerZhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao et al.ICCV 2023 · 56 citations
