Self-Supervised Human Depth Estimation From Monocular Videos
Feitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu, Marc Pollefeys, Ping Tan
Abstract
Previous methods on estimating detailed human depth often require supervised training with 'ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of the learned network. The self-supervised learning is achieved by minimizing a photo-consistency loss, which is evaluated between a video frame and its neighboring frames warped according to the estimated depth and the 3D non-rigid motion of the human body. To solve this non-rigid motion, we first estimate a rough SMPL model at each video frame and compute the non-rigid body motion accordingly, which enables self-supervised learning on estimating the shape details. Experiments demonstrate that our method enjoys better generalization and performs much better on data in the wild.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c613da8-e3ac-468b-a8e5-6d24042a1b84Cited by top-tier papers6
- Registering Explicit to Implicit: Towards High-Fidelity Garment mesh Reconstruction from Single ImagesHeming Zhu, Lingteng Qiu, Yuda Qiu, Xiaoguang HanCVPR 2022 · 32 citations
- RAFaRe: Learning Robust and Accurate Non-parametric 3D Face Reconstruction from Pseudo 2D&3D PairsLongwei Guo, Hao Zhu, Yuanxun Lu, Menghua Wu et al.AAAI 2023 · 14 citations
- Fusing the Old with the New: Learning Relative Camera Pose with Geometry-Guided UncertaintyBingbing Zhuang, Manmohan ChandrakerCVPR 2021
- Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance VideosYasamin Jafarian, Hyun Soo ParkCVPR 2021
- Recurrent Multi-View Alignment Network for Unsupervised Surface RegistrationWanquan Feng, Juyong Zhang, Hongrui Cai, Haofei Xu et al.CVPR 2021
Builds on6
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai et al.ICCV 2019 · 367 citations
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 343 citations
- TexturePose: Supervising Human Mesh Estimation With Texture ConsistencyGeorgios Pavlakos, Nikos Kolotouros, Kostas DaniilidisICCV 2019 · 109 citations
Related papers
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
- Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth EstimationHoang Chuong Nguyen, Tianyu Wang, José M. Álvarez, Miaomiao LiuCVPR 2024 · 5 citations
- iVS-Net: Learning Human View Synthesis from Internet VideosJunting Dong, Qi Fang, Tianshuo Yang, Qing Shuai et al.ICCV 2023 · 9 citations
- AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose OptimizationKun Wang, Zhiqiang Yan, Huang Tian, Zhenyu Zhang et al.AAAI 2024 · 6 citations
- Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the WildKaifeng Zhang, Yang Fu, Shubhankar Borse, Hong Cai et al.ICLR 2023 · 8 citations
