Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos
Yasamin Jafarian, Hyun Soo Park
Abstract
A key challenge of learning a visual representation for the 3D high fidelity geometry of dressed humans lies in the limited availability of the ground truth data (e.g., 3D scanned models), which results in the performance degradation of 3D human reconstruction when applying to real-world imagery. We address this challenge by leveraging a new data resource: a number of social media dance videos that span diverse appearance, clothing styles, performances, and identities. Each video depicts dynamic movements of the body and clothes of a single person while lacking the 3D ground truth geometry. To learn a visual representation from these videos, we present a new self-supervised learning method to use the local transformation that warps the predicted local geometry of the person from an image to that of another image at a different time instant. This allows self-supervision by enforcing a temporal coherence over the predictions. In addition, we jointly learn the depths along with the surface normals that are highly responsive to local texture, wrinkle, and shade by maximizing their geometric consistency. Our method is end-to-end trainable, resulting in high fidelity depth estimation that predicts fine geometry faithful to the input real image. We further provide a theoretical bound of self-supervised learning via an uncertainty analysis that characterizes the performance of the self-supervised learning without training. We demonstrate that our method outperforms the state-of-the-art human depth estimation and human shape recovery approaches on both real and rendered images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49e94872-b892-423c-bca1-ef85ff6212aaCited by top-tier papers66
- ICON: Implicit Clothed humans Obtained from NormalsYuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. BlackCVPR 2022 · 286 citations
- MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware DiffusionDi Chang, Yichun Shi, Quankai Gao, Hongyi Xu et al.ICML 2024 · 125 citations
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan et al.CVPR 2022 · 113 citations
- MagicAnimate: Temporally Consistent Human Image Animation using Diffusion ModelZhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan et al.CVPR 2024 · 106 citations
- Stable Video Infinity: Infinite-Length Video Generation with Error RecyclingWuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao et al.ICLR 2026 · 69 citations
Builds on11
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai et al.ICCV 2019 · 367 citations
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 343 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- A Neural Network for Detailed Human Depth Estimation From a Single ImageSicong Tang, Feitong Tan, Kelvin Cheng, Zhaoyang Li et al.ICCV 2019 · 46 citations
Related papers
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu et al.CVPR 2020
- iVS-Net: Learning Human View Synthesis from Internet VideosJunting Dong, Qi Fang, Tianshuo Yang, Qing Shuai et al.ICCV 2023 · 9 citations
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 142 citations
- Learning Realistic Human Reposing using Cyclic Self-Supervision with 3D Shape, Pose, and Appearance ConsistencySoubhik Sanyal, Betty J. Mohler, Alex Vorobiov, Larry Davis et al.ICCV 2021 · 20 citations
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
