Multi-Person Implicit Reconstruction From a Single Image
Armin Mustafa, Akin Caliskan, Lourdes Agapito, Adrian Hilton
Abstract
We present a new end-to-end learning framework to obtain detailed and spatially coherent reconstructions of multiple people from a single image. Existing multi-person methods suffer from two main drawbacks: they are often model-based and therefore cannot capture accurate 3D models of people with loose clothing and hair; or they require manual intervention to resolve occlusions or interactions. Our method addresses both limitations by introducing the first end-to-end learning approach to perform modelfree implicit reconstruction for realistic 3D capture of multiple clothed people in arbitrary poses (with occlusions) from a single image. Our network simultaneously estimates the 3D geometry of each person and their 6DOF spatial locations, to obtain a coherent multi-human reconstruction. In addition, we introduce a new synthetic dataset that depicts images with a varying number of inter-occluded humans and a variety of clothing and hair styles. We demonstrate robust, high-resolution reconstructions on images of multiple humans with complex occlusions, loose clothing and a large variety of poses and scenes. Our quantitative evaluation on both synthetic and real world datasets demonstrates state-of-the-art performance with significant improvements in the accuracy and completeness of the reconstructions over competing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Photorealistic Monocular 3D Reconstruction of Humans Wearing ClothingThiemo Alldieck, Mihai Zanfir, Cristian SminchisescuCVPR 2022 · 136 citations
- Human Mesh Recovery from Multiple ShotsGeorgios Pavlakos, Jitendra Malik, Angjoo KanazawaCVPR 2022 · 42 citations
- Reconstructing Groups of People with Hypergraph Relational ReasoningBuzhen Huang, Jingyi Ju, Zhihao Li, Yangang WangICCV 2023 · 21 citations
- ContactField: Implicit Field Representation for Multi-Person Interaction GeometryHansol Lee, Tackgeun You, Hansoo Park, Woohyeon Shim et al.NeurIPS 2024 · 2 citations
- Human Interaction-Aware 3D Reconstruction from a Single ImageGwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park et al.CVPR 2026
Builds on17
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai et al.ICCV 2019 · 367 citations
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 343 citations
Related papers
- DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided DecouplingYiming Jiang, Wenfeng Song, Shuai Li, Aimin HaoAAAI 2026
- High-Fidelity Clothed Avatar Reconstruction from a Single ImageTingting Liao, Xiaomei Zhang, Yuliang Xiu, Hongwei Yi et al.CVPR 2023
- Monocular, One-stage, Regression of Multiple 3D PeopleYu Sun, Qian Bao, Wu Liu, Yili Fu et al.ICCV 2021 · 335 citations
- AG3D: Learning to Generate 3D Avatars from 2D Image CollectionsZijian Dong, Xu Chen, Jinlong Yang, Michael J. Black et al.ICCV 2023 · 76 citations
- CoMotion: Concurrent Multi-person 3D MotionAlejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter et al.ICLR 2025
