Authentic volumetric avatars from a phone scan
Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhöfer, Shunsuke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, Jason M. Saragih
Abstract
Creating photorealistic avatars of existing people currently requires extensive person-specific data capture, which is usually only accessible to the VFX industry and not the general public. Our work aims to address this drawback by relying only on a short mobile phone capture to obtain a drivable 3D head avatar that matches a person's likeness faithfully. In contrast to existing approaches, our architecture avoids the complex task of directly modeling the entire manifold of human appearance, aiming instead to generate an avatar model that can be specialized to novel identities using only small amounts of data. The model dispenses with low-dimensional latent spaces that are commonly employed for hallucinating novel identities, and instead, uses a conditional representation that can extract person-specific information at multiple scales from a high resolution registered neutral phone scan. We achieve high quality results through the use of a novel universal avatar prior that has been trained on high resolution multi-view video captures of facial performances of hundreds of human subjects. By fine-tuning the model using inverse rendering we achieve increased realism and personalize its range of motion. The output of our approach is not only a high-fidelity 3D head avatar that matches the person's facial shape and appearance, but one that can also be driven using a jointly discovered shared global expression space with disentangled controls for gaze direction. Via a series of experiments we demonstrate that our avatars are faithful representations of the subject's likeness. Compared to other state-of-the-art methods for lightweight avatar creation, our approach exhibits superior visual quality and animateability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 64c9c5b2-c17b-4b40-ac4d-9123448756b0Cited by top-tier papers61
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
- Encoder-based Domain Tuning for Fast Personalization of Text-to-Image ModelsRinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano et al.SIGGRAPH 2023 · 154 citations
- Gaussian Head Avatar: Ultra High-Fidelity Head Avatar via Dynamic GaussiansYuelang Xu, Bengwang Chen, Zhe Li, Hongwen Zhang et al.CVPR 2024 · 84 citations
- AvatarReX: Real-time Expressive Full-body AvatarsZerong Zheng, Xiaochen Zhao, Hongwen Zhang, Boning Liu et al.SIGGRAPH 2023 · 80 citations
- Neural Localizer Fields for Continuous 3D Human Pose and Shape EstimationIstván Sárándi, Gerard Pons-MollNeurIPS 2024 · 76 citations
Related papers
- Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorChen Guo, Junxuan Li, Yash Kant, Yaser Sheikh et al.CVPR 2025
- HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsByungjun Kim, Shunsuke Saito, Giljoo Nam, Tomas Simon et al.ICCV 2025 · 2 citations
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong et al.CVPR 2025
- Avat3r: Large Animatable Gaussian Reconstruction Model for High-Fidelity 3D Head AvatarsTobias Kirschstein, Javier Romero, Artem Sevastopolsky, Matthias Nießner et al.ICCV 2025 · 10 citations
- Pixel-Aligned Volumetric AvatarsAmit Raj, Michael Zollhöfer, Tomas Simon, Jason M. Saragih et al.CVPR 2021
