Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos
Yasamin Jafarian, Hyun Soo Park
摘要
A key challenge of learning a visual representation for the 3D high fidelity geometry of dressed humans lies in the limited availability of the ground truth data (e.g., 3D scanned models), which results in the performance degradation of 3D human reconstruction when applying to real-world imagery. We address this challenge by leveraging a new data resource: a number of social media dance videos that span diverse appearance, clothing styles, performances, and identities. Each video depicts dynamic movements of the body and clothes of a single person while lacking the 3D ground truth geometry. To learn a visual representation from these videos, we present a new self-supervised learning method to use the local transformation that warps the predicted local geometry of the person from an image to that of another image at a different time instant. This allows self-supervision by enforcing a temporal coherence over the predictions. In addition, we jointly learn the depths along with the surface normals that are highly responsive to local texture, wrinkle, and shade by maximizing their geometric consistency. Our method is end-to-end trainable, resulting in high fidelity depth estimation that predicts fine geometry faithful to the input real image. We further provide a theoretical bound of self-supervised learning via an uncertainty analysis that characterizes the performance of the self-supervised learning without training. We demonstrate that our method outperforms the state-of-the-art human depth estimation and human shape recovery approaches on both real and rendered images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper66
- ICON: Implicit Clothed humans Obtained from NormalsYuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. BlackCVPR 2022 · 被引用 286 次
- MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware DiffusionDi Chang, Yichun Shi, Quankai Gao, Hongyi Xu 等ICML 2024 · 被引用 125 次
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan 等CVPR 2022 · 被引用 113 次
- MagicAnimate: Temporally Consistent Human Image Animation using Diffusion ModelZhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan 等CVPR 2024 · 被引用 106 次
- Stable Video Infinity: Infinite-Length Video Generation with Error RecyclingWuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao 等ICLR 2026 · 被引用 69 次
它引用的顶会 Paper11
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai 等ICCV 2019 · 被引用 367 次
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 被引用 343 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- A Neural Network for Detailed Human Depth Estimation From a Single ImageSicong Tang, Feitong Tan, Kelvin Cheng, Zhaoyang Li 等ICCV 2019 · 被引用 46 次
相关 Paper
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu 等CVPR 2020
- iVS-Net: Learning Human View Synthesis from Internet VideosJunting Dong, Qi Fang, Tianshuo Yang, Qing Shuai 等ICCV 2023 · 被引用 9 次
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 被引用 142 次
- Learning Realistic Human Reposing using Cyclic Self-Supervision with 3D Shape, Pose, and Appearance ConsistencySoubhik Sanyal, Betty J. Mohler, Alex Vorobiov, Larry Davis 等ICCV 2021 · 被引用 20 次
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 被引用 1 次
