REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak Supervision
Mihai Fieraru, Mihai Zanfir, Teodor Alexandru Szente, Eduard Gabriel Bazavan, Vlad Olaru, Cristian Sminchisescu
摘要
The three-dimensional reconstruction of multiple interacting humans given a monocular image is crucial for the general task of scene understanding, as capturing the subtleties of interaction is often the very reason for taking a picture. Current 3D human reconstruction methods either treat each person independently, ignoring most of the context, or reconstruct people jointly, but cannot recover interactions correctly when people are in close proximity. In this work, we introduce REMIPS, a model for 3D Reconstruction of Multiple Interacting People under Weak Supervision. REMIPS can reconstruct a variable number of people directly from monocular images. At the core of our methodology stands a novel transformer network that combines unordered person tokens (one for each detected human) with positional-encoded tokens from image features patches. We introduce a novel unified model for self-and interpenetration-collisions based on a mesh approximation computed by applying decimation operators. We rely on self-supervised losses for flexibility and generalisation in-the-wild and incorporate self-contact and interaction-contact losses directly into the learning process. With REMIPS, we report state-of-the-art quantitative results on common benchmarks even in cases where no 3D supervision is used. Additionally, qualitative visual results show that our reconstructions are plausible in terms of pose and shape and coherent for challenging images, collected in-the-wild, where people are often interacting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- DECO: Dense Estimation of 3D Human-Scene Contact In The WildShashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi 等ICCV 2023 · 被引用 54 次
- Differentiable Dynamics for Articulated 3d Human Motion ReconstructionErik Gärtner, Mykhaylo Andriluka, Erwin Coumans, Cristian SminchisescuCVPR 2022 · 被引用 33 次
- JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh RecoveryJiahao Li, Zongxin Yang, Xiaohan Wang, Jianxin Ma 等ICCV 2023 · 被引用 22 次
- Generative Proxemics: A Prior for 3D Social Interaction from ImagesLea Müller, Vickie Ye, Georgios Pavlakos, Michael J. Black 等CVPR 2024 · 被引用 14 次
- DPMesh: Exploiting Diffusion Prior for Occluded Human Mesh RecoveryYixuan Zhu, Ao Li, Yansong Tang, Wenliang Zhao 等CVPR 2024 · 被引用 10 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 被引用 384 次
- THUNDR: Transformer-based 3D HUmaN Reconstruction with MarkersMihai Zanfir, Andrei Zanfir, Eduard Gabriel Bazavan, William T. Freeman 等ICCV 2021 · 被引用 75 次
- Learning Complex 3D Human Self-ContactMihai Fieraru, Mihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa 等AAAI 2021 · 被引用 44 次
相关 Paper
- MultiPly: Reconstruction of Multiple People from Monocular Video in the WildZeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang 等CVPR 2024
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 被引用 1 次
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri 等CVPR 2025
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan 等CVPR 2024
