Reconstructing Groups of People with Hypergraph Relational Reasoning
Buzhen Huang, Jingyi Ju, Zhihao Li, Yangang Wang
Abstract
Due to the mutual occlusion, severe scale variation, and complex spatial distribution, the current multi-person mesh recovery methods cannot produce accurate absolute body poses and shapes in large-scale crowded scenes. To address the obstacles, we fully exploit crowd features for reconstructing groups of people from a monocular image. A novel hypergraph relational reasoning network is proposed to formulate the complex and high-order relation correlations among individuals and groups in the crowd. We first extract compact human features and location information from the original high-resolution image. By conducting the relational reasoning on the extracted individual features, the underlying crowd collectiveness and interaction relationship can provide additional group information for the reconstruction. Finally, the updated individual features and the localization information are used to regress human meshes in camera coordinates. To facilitate the network training, we further build pseudo ground-truth on two crowd datasets, which may also promote future research on pose estimation and human behavior understanding in crowded scenes. The experimental results show that our approach outperforms other baseline methods both in crowded and common scenarios. The code and datasets are publicly available at https://github.com/boycehbz/GroupRec.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1a05dcc-0a92-4034-8672-b7e9da22bc58Cited by top-tier papers6
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 6 citations
- TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view VideosJinpeng Liu, Yukang Xu, Yutong Li, Xingyu LiuCVPR 2026 · 1 citation
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan et al.CVPR 2024
- Crowd4D: Scene-Aware Monocular 4D Crowd ReconstructionHongbo Kang, Tianyi Zhou, Qingyang Yang, Hongwei wen et al.ICML 2026
- MultiPly: Reconstruction of Multiple People from Monocular Video in the WildZeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang et al.CVPR 2024
Builds on28
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 368 citations
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu et al.SIGGRAPH 2020 · 267 citations
- EvolveGraph: Multi-Agent Trajectory Prediction with Dynamic Relational ReasoningJiachen Li, Fan Yang, Masayoshi Tomizuka, Chiho ChoiNeurIPS 2020 · 258 citations
Related papers
- Learning to Estimate Robust 3D Human Mesh from In-the-Wild Crowded ScenesHongsuk Choi, Gyeongsik Moon, JoonKyu Park, Kyoung Mu LeeCVPR 2022 · 92 citations
- Crowd3D: Towards Hundreds of People Reconstruction from a Single ImageHao Wen, Jing Huang, Huili Cui, Haozhe Lin et al.CVPR 2023
- MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular ImagesChentao Song, He Zhang, Haolei Yuan, Haozhe Lin et al.CVPR 2026 · 5 citations
- Dynamic Graph Reasoning for Multi-person 3D Pose EstimationZhongwei Qiu, Qiansheng Yang, Jian Wang, Dongmei FuACM MM 2022 · 14 citations
- Learning Human Mesh Recovery in 3D ScenesZehong Shen, Zhi Cen, Sida Peng, Qing Shuai et al.CVPR 2023
