Learning to Estimate Robust 3D Human Mesh from In-the-Wild Crowded Scenes
Hongsuk Choi, Gyeongsik Moon, JoonKyu Park, Kyoung Mu Lee
Abstract
We consider the problem of recovering a single person's 3D human mesh from in-the-wild crowded scenes. While much progress has been in 3D human mesh estimation, existing methods struggle when test input has crowded scenes. The first reason for the failure is a domain gap between training and testing data. A motion capture dataset, which provides accurate 3D labels for training, lacks crowd data and impedes a network from learning crowded scene-robust image features of a target person. The second reason is a feature processing that spatially averages the feature map of a localized bounding box containing multiple people. Averaging the whole feature map makes a target person's feature indistinguishable from others. We present 3DCrowdNet that firstly explicitly targets in-the-wild crowded scenes and estimates a robust 3D human mesh by addressing the above issues. First, we leverage 2D human pose estimation that does not require a motion capture dataset with 3D labels for training and does not suffer from the domain gap. Second, we propose a joint-based regressor that distinguishes a target person's feature from others. Our joint-based regressor preserves the spatial activation of a target by sampling features from the target's joint locations and regresses human model parameters. As a result, 3DCrowdNet learns target-focused features and effectively excludes the irrelevant features of nearby persons. We conduct experiments on various benchmarks and prove the robustness of 3D CrowdNet to the in-the-wild crowded scenes both quantitatively and qualitatively. Codes are available here <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/hongsukchoi/3DCrowdNet_RELEASE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7369812-ba1a-4b55-a3fe-b988456d0a4fCited by top-tier papers38
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Putting People in their Place: Monocular Regression of 3D People in DepthYu Sun, Wu Liu, Qian Bao, Yili Fu et al.CVPR 2022 · 152 citations
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi et al.CVPR 2022 · 116 citations
- HAP: Structure-Aware Masked Image Modeling for Human-Centric PerceptionJunkun Yuan, Xinyu Zhang, Hao Zhou, Jian Wang et al.NeurIPS 2023 · 46 citations
- Distribution-Aligned Diffusion for Human Mesh RecoveryLin Geng Foo, Jia Gong, Hossein Rahmani, Jun LiuICCV 2023 · 37 citations
Builds on8
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 368 citations
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu et al.SIGGRAPH 2020 · 267 citations
- Coherent Reconstruction of Multiple Humans From a Single ImageWen Jiang, Nikos Kolotouros, Georgios Pavlakos, Xiaowei Zhou et al.CVPR 2020
- Reconstructing 3D Human Pose by Watching Humans in the MirrorQi Fang, Qing Shuai, Junting Dong, Hujun Bao et al.CVPR 2021
Related papers
- Reconstructing Groups of People with Hypergraph Relational ReasoningBuzhen Huang, Jingyi Ju, Zhihao Li, Yangang WangICCV 2023 · 21 citations
- Instance-Aware Contrastive Learning for Occluded Human Mesh ReconstructionMi-Gyeong Gwon, Gi-Mun Um, Won-Sik Cheong, Wonjun KimCVPR 2024
- Progressive Multi-View Human Mesh Recovery with Self-SupervisionXuan Gong, Liangchen Song, Meng Zheng, Benjamin Planche et al.AAAI 2023 · 16 citations
- DC-GNet: Deep Mesh Relation Capturing Graph Convolution Network for 3D Human Shape ReconstructionShihao Zhou, Mengxi Jiang, Shanshan Cai, Yunqi LeiACM MM 2021 · 14 citations
- DecenterNet: Bottom-Up Human Pose Estimation Via Decentralized Pose RepresentationTao Wang, Lei Jin, Zhang Wang, Xiaojin Fan et al.ACM MM 2023 · 14 citations
