ICON: Implicit Clothed humans Obtained from Normals
Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. Black
Abstract
Current methods for learning realistic and animatable 3D clothed avatars need either posed 3D scans or 2D images with carefully controlled user poses. In contrast, our goal is to learn an avatar from only 2D images of people in unconstrained poses. Given a set of images, our method estimates a detailed 3D surface from each image and then combines these into an animatable avatar. Implicit functions are well suited to the first task, as they can capture details like hair and clothes. Current methods, however, are not robust to varied human poses and often produce 3D surfaces with broken or disembodied limbs, missing details, or non-human shapes. The problem is that these methods use global feature encoders that are sensitive to global pose. To address this, we propose ICON (“Tmplicit Clothed humans Obtained from Normals”), which, instead, uses local features. ICON has two main modules, both of which exploit the SMPL(-X) body model. First, ICON infers detailed clothed-human normals (front/back) conditioned on the SMPL(-X) normals. Second, a visibility-aware implicit surface regressor produces an iso-surface of a human occupancy field. Importantly, at inference time, a feedback loop alternates between refining the SMPL(-X) mesh using the inferred clothed normals and then refining the normals. Given multiple reconstructed frames of a subject in varied poses, we use a modified version of SCANimate to produce an animatable avatar from them. Evaluation on the AGORA and CAPE datasets shows that ICON outperforms the state of the art in reconstruction, even with heavily limited training data. Additionally, it is much more robust to out-of-distribution samples, e.g., in-the-wild poses/images and out-of-frame cropping. ICON takes a step towards robust 3D clothed human reconstruction from in-the-wild images. This enables avatar creation directly from video with personalized pose-dependent cloth deformation. Models and code are available for research at https://icon.is.tue.mpg.de.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0db46a4-bf10-4025-a5d3-33b927bc0bb5Cited by top-tier papers154
- Putting People in their Place: Monocular Regression of 3D People in DepthYu Sun, Wu Liu, Qian Bao, Yili Fu et al.CVPR 2022 · 152 citations
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 142 citations
- SHERF: Generalizable Human NeRF from a Single ImageShoukang Hu, Fangzhou Hong, Liang Pan, Haiyi Mei et al.ICCV 2023 · 115 citations
- DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric RenderingWei Cheng, Ruixiang Chen, Siming Fan, Wanqi Yin et al.ICCV 2023 · 106 citations
- GPS-Gaussian: Generalizable Pixel-Wise 3D Gaussian Splatting for Real-Time Human Novel View SynthesisShunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu et al.CVPR 2024 · 79 citations
Builds on31
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopHongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang et al.ICCV 2021 · 376 citations
Related papers
- SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar NetworksShunsuke Saito, Jinlong Yang, Qianli Ma, Michael J. BlackCVPR 2021
- ECON: Explicit Clothed humans Optimized via Normal integrationYuliang Xiu, Jinlong Yang, Xu Cao, Dimitrios Tzionas et al.CVPR 2023
- ARCH: Animatable Reconstruction of Clothed HumansZeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li et al.CVPR 2020
- SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human ReconstructionZechuan Zhang, Zongxin Yang, Yi YangCVPR 2024 · 44 citations
- Neural-GIF: Neural Generalized Implicit Functions for Animating People in ClothingGarvita Tiwari, Nikolaos Sarafianos, Tony Tung, Gerard Pons-MollICCV 2021 · 130 citations
