Putting People in their Place: Monocular Regression of 3D People in Depth
Yu Sun, Wu Liu, Qian Bao, Yili Fu, Tao Mei, Michael J. Black
Abstract
Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene contains people of very different sizes, e.g. from infants to adults. To solve this, we need several things. First, we develop a novel method to infer the poses and depth of multiple people in a single image. While previous work that estimates multiple people does so by reasoning in the image plane, our method, called BEV, adds an additional imaginary Bird's-Eye-View representation to explicitly reason about depth. BEV reasons simultaneously about body centers in the image and in depth and, by combing these, estimates 3D body position. Unlike prior work, BEV is a single-shot method that is end-to-end differentiable. Second, height varies with age, making it impossible to resolve depth without also estimating the age of people in the image. To do so, we exploit a 3D body model space that lets BEV infer shapes from infants to adults. Third, to train BEV, we need a new dataset. Specifically, we create a "Relative Human" (RH) dataset that includes age labels and relative depth relationships between the people in the images. Extensive experiments on RH and AGORA demonstrate the effectiveness of the model and training scheme. BEV outperforms existing methods on depth reasoning, child shape estimation, and robustness to occlusion. The code and dataset are released for research purposes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a12a9842-8cbd-4b54-9121-e2189097764eCited by top-tier papers77
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- ICON: Implicit Clothed humans Obtained from NormalsYuliang Xiu, Jinlong Yang, Dimitrios Tzionas, Michael J. BlackCVPR 2022 · 286 citations
- Gait Recognition in the Wild with Dense 3D Representations and A BenchmarkJinkai Zheng, Xinchen Liu, Wu Liu, Lingxiao He et al.CVPR 2022 · 228 citations
- Capturing and Inferring Dense Full-Body Human-Scene ContactChun-Hao P. Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin et al.CVPR 2022 · 106 citations
- EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the WildManuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen et al.ICCV 2023 · 94 citations
Builds on20
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 398 citations
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopHongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang et al.ICCV 2021 · 376 citations
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 368 citations
Related papers
- AGORA: Avatars in Geography Optimized for Regression AnalysisPriyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann et al.CVPR 2021
- From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera CalibrationZekun Qian, Ruize Han, Wei Feng, Song WangCVPR 2024 · 8 citations
- Mutual Adaptive Reasoning for Monocular 3D Multi-Person Pose EstimationJuze Zhang, Jingya Wang, Ye Shi, Fei Gao et al.ACM MM 2022 · 15 citations
- Accurate Estimation of Body Height From a Single Depth Image via a Four-Stage Developing NetworkFukun Yin, Shizhe ZhouCVPR 2020
- Monocular, One-stage, Regression of Multiple 3D PeopleYu Sun, Qian Bao, Wu Liu, Yili Fu et al.ICCV 2021 · 335 citations
