Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
Simon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito, Matthias Nießner
摘要
We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a 3D morphable face model (3DMM). We exploit the latent features of the DINO foundation model, and introduce a tailored surface normal and uv-coordinate prediction head. We train our model by registering three high-quality 3D face datasets against the FLAME mesh topology, which results in a total of over 1,000 identities and 976K images. For 3D face reconstruction, we propose a FLAME fitting opitmization that solves for the 3DMM parameters from the uv-coordinate and normal estimates. To evaluate our method, we introduce a new benchmark for single-image face reconstruction, which features high diversity facial expressions, viewing angles, and ethnicities. Crucially, our benchmark is the first to evaluate both posed and neutral facial geometry. Ultimately, our method outperforms the most competitive baselines by over 15% in terms of geometric accuracy for posed facial expressions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionTobias Kirschstein, Simon Giebenhain, Matthias NießnerCVPR 2026 · 被引用 10 次
- UIKA: Fast Universal Head Avatar from Pose-Free ImagesZijian Wu, Boyao Zhou, Liangxiao Hu, Hongyu Liu 等CVPR 2026 · 被引用 6 次
- FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationXinya Ji, Sebastian Weiss, Manuel Kansy, Jacek Naruniec 等ICLR 2026 · 被引用 6 次
- FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed DeformationCheng Peng, Zhuo Su, Liao Wang, Chen Guo 等CVPR 2026 · 被引用 2 次
- VGGTFace: Topologically Consistent Facial Geometry Reconstruction in the WildXin Ming, Yuxuan Han, Tianyu Huang, Feng XuAAAI 2026 · 被引用 2 次
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 被引用 662 次
相关 Paper
- Towards High-Fidelity 3D Face Reconstruction From In-the-Wild Images Using Graph Convolutional NetworksJiangke Lin, Yi Yuan, Tianjia Shao, Kun ZhouCVPR 2020
- WarpHE4D: Dense 4D Head Map Toward Full Head ReconstructionJongseob Yun, Yong-Hoon Kwon, Min-Gyu Park, Ju-Mi Kang 等ICCV 2025 · 被引用 1 次
- Reconstructing Humans with a Biomechanically Accurate SkeletonYan Xia, Xiaowei Zhou, Etienne Vouga, Qixing Huang 等CVPR 2025
- Relightify: Relightable 3D Faces from a Single Image via Diffusion ModelsFoivos Paraperas Papantoniou, Alexandros Lattas, Stylianos Moschoglou, Stefanos ZafeiriouICCV 2023 · 被引用 40 次
- Densemarks: Learning Canonical Embeddings for Human Heads Images via Point TracksDmitrii Pozdeev, Alexey Artemov, Ananta R. Bhattarai, Artem SevastopolskyICLR 2026
