REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak Supervision
Mihai Fieraru, Mihai Zanfir, Teodor Alexandru Szente, Eduard Gabriel Bazavan, Vlad Olaru, Cristian Sminchisescu
Abstract
The three-dimensional reconstruction of multiple interacting humans given a monocular image is crucial for the general task of scene understanding, as capturing the subtleties of interaction is often the very reason for taking a picture. Current 3D human reconstruction methods either treat each person independently, ignoring most of the context, or reconstruct people jointly, but cannot recover interactions correctly when people are in close proximity. In this work, we introduce REMIPS, a model for 3D Reconstruction of Multiple Interacting People under Weak Supervision. REMIPS can reconstruct a variable number of people directly from monocular images. At the core of our methodology stands a novel transformer network that combines unordered person tokens (one for each detected human) with positional-encoded tokens from image features patches. We introduce a novel unified model for self-and interpenetration-collisions based on a mesh approximation computed by applying decimation operators. We rely on self-supervised losses for flexibility and generalisation in-the-wild and incorporate self-contact and interaction-contact losses directly into the learning process. With REMIPS, we report state-of-the-art quantitative results on common benchmarks even in cases where no 3D supervision is used. Additionally, qualitative visual results show that our reconstructions are plausible in terms of pose and shape and coherent for challenging images, collected in-the-wild, where people are often interacting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f46313-0dfc-4eeb-8cbf-ddaa08ec4a59Cited by top-tier papers15
- DECO: Dense Estimation of 3D Human-Scene Contact In The WildShashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi et al.ICCV 2023 · 54 citations
- Differentiable Dynamics for Articulated 3d Human Motion ReconstructionErik Gärtner, Mykhaylo Andriluka, Erwin Coumans, Cristian SminchisescuCVPR 2022 · 33 citations
- JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh RecoveryJiahao Li, Zongxin Yang, Xiaohan Wang, Jianxin Ma et al.ICCV 2023 · 22 citations
- Generative Proxemics: A Prior for 3D Social Interaction from ImagesLea Müller, Vickie Ye, Georgios Pavlakos, Michael J. Black et al.CVPR 2024 · 14 citations
- DPMesh: Exploiting Diffusion Prior for Occluded Human Mesh RecoveryYixuan Zhu, Ao Li, Yansong Tang, Wenliang Zhao et al.CVPR 2024 · 10 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- THUNDR: Transformer-based 3D HUmaN Reconstruction with MarkersMihai Zanfir, Andrei Zanfir, Eduard Gabriel Bazavan, William T. Freeman et al.ICCV 2021 · 75 citations
- Learning Complex 3D Human Self-ContactMihai Fieraru, Mihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa et al.AAAI 2021 · 44 citations
Related papers
- MultiPly: Reconstruction of Multiple People from Monocular Video in the WildZeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang et al.CVPR 2024
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan et al.CVPR 2024
