PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing
Antonio Oroz, Matthias Nießner, Tobias Kirschstein
Abstract
We present PercHead, a model for single-image 3D head reconstruction and disentangled 3D editing - two tasks that are inherently challenging due to ambiguity in plausible explanations for the same input. At the heart of our approach lies our novel perceptual loss based on DINOv2 and SAM 2.1. Unlike widely-adopted low-level losses like LPIPS, SSIM or L1, we rely on deep visual understanding of images and the resulting generalized supervision signals. We show that our new loss can be a drop-in replacement for standard losses and used to improve visual quality in high-frequency areas. We base our model architecture on Vision Transformers (ViTs), allowing us to decouple the 3D representation from the 2D input. We train our method on multi-view images for view-consistency and in-the-wild images for strong transferability to new environments. Our model achieves state-of-the-art performance in novel-view synthesis and, furthermore, exhibits exceptional robustness to extreme viewing angles. We also extend our base model to disentangled 3D editing by swapping the encoder and fine-tuning the network. A segmentation map controls geometry and either a text prompt or a reference image specifies appearance. We highlight the intuitive and powerful 3D editing capabilities through an interactive GUI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b437b1df-6c3c-4a08-82bb-0106f36fe8dbCited by top-tier papers3
- FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionTobias Kirschstein, Simon Giebenhain, Matthias NießnerCVPR 2026 · 10 citations
- UIKA: Fast Universal Head Avatar from Pose-Free ImagesZijian Wu, Boyao Zhou, Liangxiao Hu, Hongyu Liu et al.CVPR 2026 · 6 citations
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view GenerationAviral Chharia, Fernando De la TorreCVPR 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- PointDiT: Pixel-Space Diffusion for Monocular Geometry EstimationHaofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek et al.ICML 2026 · 3 citations
- Positional Encoding FieldYunpeng Bai, Haoxiang Li, Qixing HuangICLR 2026 · 22 citations
- ViTok-v2: Scaling Native Resolution Autoencoders to 5 Billion ParametersPhilippe Hansen-Estruch, Jiahui Chen, Vivek Ramanujan, Orr Zohar et al.ICML 2026
- One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step DiffusionYitong Dong, Qi Zhang, Minchao Jiang, Zhiqiang Wu et al.AAAI 2026 · 2 citations
- REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak SupervisionMihai Fieraru, Mihai Zanfir, Teodor Alexandru Szente, Eduard Gabriel Bazavan et al.NeurIPS 2021 · 44 citations
