Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
Simon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito, Matthias Nießner
Abstract
We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a 3D morphable face model (3DMM). We exploit the latent features of the DINO foundation model, and introduce a tailored surface normal and uv-coordinate prediction head. We train our model by registering three high-quality 3D face datasets against the FLAME mesh topology, which results in a total of over 1,000 identities and 976K images. For 3D face reconstruction, we propose a FLAME fitting opitmization that solves for the 3DMM parameters from the uv-coordinate and normal estimates. To evaluate our method, we introduce a new benchmark for single-image face reconstruction, which features high diversity facial expressions, viewing angles, and ethnicities. Crucially, our benchmark is the first to evaluate both posed and neutral facial geometry. Ultimately, our method outperforms the most competitive baselines by over 15% in terms of geometric accuracy for posed facial expressions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63fa234f-935b-4741-ba88-d566d20552f3Cited by top-tier papers14
- FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionTobias Kirschstein, Simon Giebenhain, Matthias NießnerCVPR 2026 · 10 citations
- UIKA: Fast Universal Head Avatar from Pose-Free ImagesZijian Wu, Boyao Zhou, Liangxiao Hu, Hongyu Liu et al.CVPR 2026 · 6 citations
- FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationXinya Ji, Sebastian Weiss, Manuel Kansy, Jacek Naruniec et al.ICLR 2026 · 6 citations
- FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed DeformationCheng Peng, Zhuo Su, Liao Wang, Chen Guo et al.CVPR 2026 · 2 citations
- VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the WildXin Ming, Yuxuan Han, Tianyu Huang, Feng XuAAAI 2026 · 2 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
Related papers
- Towards High-Fidelity 3D Face Reconstruction From In-the-Wild Images Using Graph Convolutional NetworksJiangke Lin, Yi Yuan, Tianjia Shao, Kun ZhouCVPR 2020
- WarpHE4D: Dense 4D Head Map Toward Full Head ReconstructionJongseob Yun, Yong-Hoon Kwon, Min-Gyu Park, Ju-Mi Kang et al.ICCV 2025 · 1 citation
- Reconstructing Humans with a Biomechanically Accurate SkeletonYan Xia, Xiaowei Zhou, Etienne Vouga, Qixing Huang et al.CVPR 2025
- Relightify: Relightable 3D Faces from a Single Image via Diffusion ModelsFoivos Paraperas Papantoniou, Alexandros Lattas, Stylianos Moschoglou, Stefanos ZafeiriouICCV 2023 · 40 citations
- Densemarks: Learning Canonical Embeddings for Human Heads Images via Point TracksDmitrii Pozdeev, Alexey Artemov, Ananta R. Bhattarai, Artem SevastopolskyICLR 2026
