Monocular, One-stage, Regression of Multiple 3D People
Yu Sun, Qian Bao, Wu Liu, Yili Fu, Michael J. Black, Tao Mei
Abstract
This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a One-stage fashion for Multiple 3D People (termed ROMP). The approach is conceptually simple, bounding box-free, and able to learn a per-pixel representation in an end-to-end manner. Our method simultaneously predicts a Body Center heatmap and a Mesh Parameter map, which can jointly describe the 3D body mesh on the pixel level. Through a body-center-guided sampling process, the body mesh parameters of all people in the image are easily extracted from the Mesh Parameter map. Equipped with such a fine-grained representation, our one-stage framework is free of the complex multi-stage process and more robust to occlusion. Compared with state-of-the-art methods, ROMP achieves superior performance on the challenging multi-person benchmarks, including 3DPW and CMU Panoptic. Experiments on crowded/occluded datasets demonstrate the robustness under various types of occlusion. The code, released at https://github.com/Arthur151/ROMP, is the first real-time implementation of monocular multi-person 3D mesh regression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- Occluded Human Mesh RecoveryRawal Khirodkar, Shashank Tripathi, Kris KitaniCVPR 2022 · 74 citations
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe et al.ICCV 2023 · 59 citations
- TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerZhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao et al.ICCV 2023 · 56 citations
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen et al.ICLR 2026 · 38 citations
- Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDARPeishan Cong, Yiteng Xu, Yiming Ren, Juze Zhang et al.AAAI 2023 · 37 citations
Builds on13
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 368 citations
- DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-CompareYuanlu Xu, Song-Chun Zhu, Tony TungICCV 2019 · 204 citations
- Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled RepresentationYu Sun, Yun Ye, Wu Liu, Wenpeng Gao et al.ICCV 2019 · 196 citations
Related papers
- Body Meshes as PointsJianfeng Zhang, Dongdong Yu, Jun Hao Liew, Xuecheng Nie et al.CVPR 2021
- Distribution-Aware Single-Stage Models for Multi-Person 3D Pose EstimationZitian Wang, Xuecheng Nie, Xiaochao Qu, Yunpeng Chen et al.CVPR 2022 · 44 citations
- Multi-Person Implicit Reconstruction From a Single ImageArmin Mustafa, Akin Caliskan, Lourdes Agapito, Adrian HiltonCVPR 2021
- Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from VideosHaoyuan Li, Haoye Dong, Hanchao Jia, Dong Huang et al.ICCV 2023 · 8 citations
- PandaNet: Anchor-Based Single-Shot Multi-Person 3D Pose EstimationAbdallah Benzine, Florian Chabot, Bertrand Luvison, Quoc Cuong Pham et al.CVPR 2020
