ANIM: Accurate Neural Implicit Model for Human Reconstruction from a Single RGB-D Image
Marco Pesavento, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier, Ziyan Wang, Chun-Han Yao, Marco Volino, Edmond Boyer, Adrian Hilton, Tony Tung
Abstract
Recent progress in human shape learning, shows that neural implicit models are effective in generating 3D hu-man surfaces from limited number of views, and even from a single RGB image. However, existing monocular approaches still struggle to recover fine geometric details such as face, hands or cloth wrinkles. They are also easily prone to depth ambiguities that result in distorted geome-tries along the camera optical axis. In this paper, we ex-plore the benefits of incorporating depth observations in the reconstruction process by introducing ANIM, a novel method that reconstructs arbitrary 3D human shapes from single-view RGB-D images with an unprecedented level of accuracy. Our model learns geometric details from both multi-resolution pixel-aligned and voxel-aligned features to leverage depth information and enable spatial relation-ships, mitigating depth ambiguities. We further enhance the quality of the reconstructed shape by introducing a depth-supervision strategy, which improves the accuracy of the signed distance field estimation of points that lie on the re-constructed surface. Experiments demonstrate that ANIM outperforms state-of-the-art works that use RGB, surface normals, point cloud or RGB-D data as input. In addition, we introduce ANIM-Real, a new multi-modal dataset comprising highquality scans paired with consumer-grade RGB-D camera, and our protocol to fine-tune ANIM, enabling highquality reconstruction from real-world human capture. https://marcopesavento.github.io/Anim/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b371d8b8-d1f6-4916-a203-142b03e2307eCited by top-tier papers5
- FIP: Endowing Robust Motion Capture on Daily Garment by Fusing Flex and Inertial SensorsRuonan Zheng, Jiawei Fang, Yuan Yao, Xiaoxia Gao et al.CHI 2025 · 5 citations
- Bézier Degradation Modeling for LiDAR-based Human Motion CaptureXiaoqi An, Lin Zhao, Jun Li, Chen Gong et al.CVPR 2026
- PGC: Physics-Based Gaussian Cloth from a Single PoseMichelle Guo, Matt Jen-Yuan Chiang, Igor Santesteban, Nikolaos Sarafianos et al.CVPR 2025
- SMVRT: Implicit Human 3D Modeling Using Sparse Multi-View Volumetric Reconstruction with Transformer FusionChuanmao Fan, Chenxi Zhao, Ye DuanCVPR 2026
- SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic PriorsYufan Wu, Xuanhong Chen, Wen Li, Shunran Jia et al.CVPR 2025
Builds on25
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai et al.ICCV 2019 · 367 citations
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 343 citations
- ICON: Implicit Clothed humans Obtained from NormalsYuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. BlackCVPR 2022 · 286 citations
- ARCH++: Animation-Ready Clothed Human Reconstruction RevisitedTong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto et al.ICCV 2021 · 233 citations
Related papers
- ARCH: Animatable Reconstruction of Clothed HumansZeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li et al.CVPR 2020
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
- TMO: Textured Mesh Acquisition of Objects with a Mobile Device by using Differentiable RenderingJaehoon Choi, Dongki Jung, Taejae Lee, Sangwook Kim et al.CVPR 2023
- Differentiable Volumetric Rendering: Learning Implicit 3D Representations Without 3D SupervisionMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerCVPR 2020
- SHERF: Generalizable Human NeRF from a Single ImageShoukang Hu, Fangzhou Hong, Liang Pan, Haiyi Mei et al.ICCV 2023 · 115 citations
