One-shot Implicit Animatable Avatars with Model-based Priors
Yangyi Huang, Hongwei Yi, Weiyang Liu, Haofan Wang, Boxi Wu, Wenxiao Wang, Binbin Lin, Debing Zhang, Deng Cai
Abstract
Existing neural rendering methods for creating human avatars typically either require dense input signals such as video or multi-view images, or leverage a learned prior from large-scale specific 3D human datasets such that reconstruction can be performed with sparse-view inputs. Most of these methods fail to achieve realistic reconstruction when only a single image is available. To enable the data-efficient creation of realistic anima table 3D humans, we propose ELICIT, a novel method for learning human-specific neural radiance fields from a single image. Inspired by the fact that humans can effortlessly estimate the body geometry and imagine full-body clothing from a single image, we leverage two priors in ELICIT: 3D geometry prior and visual semantic prior. Specifically, ELICIT utilizes the 3D body shape geometry prior from a skinned vertex-based template model (i.e., SMPL) and implements the visual clothing semantic prior with the CLIP-based pre-trained models. Both priors are used to jointly guide the optimization for creating plausible content in the invisible areas. Taking advantage of the CLIP models, ELICIT can use text descriptions to generate text-conditioned unseen regions. In order to further improve visual details, we propose a segmentation-based sampling strategy that locally refines different parts of the avatar. Comprehensive evaluations on multiple popular benchmarks, including ZJU-MoCAP, Human3.6M, and DeepFashion, show that ELICIT outperforms strong baseline methods of avatar creation when only a single image is available. The code is public for research purposes at https://huangyangyi.github.io/ELICIT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure PriorsPanwang Pan, Zhuo Su, Chenguo Lin, Zhen Fan et al.NeurIPS 2024 · 76 citations
- SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human ReconstructionZechuan Zhang, Zongxin Yang, Yi YangCVPR 2024 · 44 citations
- PoseVocab: Learning Joint-structured Pose Embeddings for Human Avatar ModelingZhe Li, Zerong Zheng, Yuxiao Liu, Boyao Zhou et al.SIGGRAPH 2023 · 34 citations
- HAVE-FUN: Human Avatar Reconstruction from Few-Shot Unconstrained ImagesXihe Yang, Xingyu Chen, Daiheng Gao, Shaohui Wang et al.CVPR 2024 · 11 citations
- DiffHuman: Probabilistic Photorealistic 3D Reconstruction of HumansAkash Sengupta, Thiemo Alldieck, Nikos Kolotouros, Enric Corona et al.CVPR 2024 · 10 citations
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- NSF: Neural Surface Fields for Human Modeling from Monocular DepthYuxuan Xue, Bharat Lal Bhatnagar, Riccardo Marin, Nikolaos Sarafianos et al.ICCV 2023 · 26 citations
- Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single ImageJooHyun Kwon, Geonhee Sim, Gyeongsik MoonCVPR 2026 · 3 citations
- Structured Local Radiance Fields for Human Avatar ModelingZerong Zheng, Han Huang, Tao Yu, Hongwen Zhang et al.CVPR 2022 · 115 citations
- ICON: Implicit Clothed humans Obtained from NormalsYuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. BlackCVPR 2022 · 286 citations
- High-Fidelity Clothed Avatar Reconstruction from a Single ImageTingting Liao, Xiaomei Zhang, Yuliang Xiu, Hongwei Yi et al.CVPR 2023
