MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild
Zeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang, Julien Valentin, Otmar Hilliges, Jie Song
Abstract
We present MultiPly a novel framework to reconstruct multiple people in 3D from monocular in-the-wild videos. Reconstructing multiple individuals moving and interacting naturally from monocular in-the-wild videos poses a challenging task. Addressing it necessitates precise pixel-level disentanglement of individuals without any prior knowledge about the subjects. Moreover it requires recovering intricate and complete 3D human shapes from short video sequences intensifying the level of difficulty. To tackle these challenges we first define a layered neural representation for the entire scene composited by individual human and background models. We learn the layered neural representation from videos via our layer-wise differentiable volume rendering. This learning process is further enhanced by our hybrid instance segmentation approach which combines the self-supervised 3D segmentation and the promptable 2D segmentation module yielding reliable instance segmentation supervision even under close human interaction. A confidence-guided optimization formulation is introduced to optimize the human poses and shape/appearance alternately. We incorporate effective objectives to refine human poses via photometric information and impose physically plausible constraints on human dynamics leading to temporally consistent 3D reconstructions with high fidelity. The evaluation of our method shows the superiority over prior art on publicly available datasets and in-the-wild videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou et al.CVPR 2026 · 12 citations
- Geo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionZeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus et al.ICCV 2025 · 9 citations
- Disentangled Clothed Avatar Generation with Layered RepresentationWeitian Zhang, Yichao Yan, Sijing Wu, Manwen Liao et al.ICCV 2025 · 3 citations
- Human Interaction-Aware 3D Reconstruction from a Single ImageGwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park et al.CVPR 2026
- GeoAvatar: Geometrically-Consistent Multi-Person Avatar Reconstruction from Sparse Multi-View VideosSoohyun Lee, Seoyeon Kim, HeeKyung Lee, Won-Sik Jeong et al.CVPR 2025
Builds on30
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 1,421 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Implicit Geometric Regularization for Learning ShapesAmos Gropp, Lior Yariv, Niv Haim, Matan Atzmon et al.ICML 2020 · 1,001 citations
Related papers
- Vid2Avatar: 3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene DecompositionChen Guo, Tianjian Jiang, Xu Chen, Jie Song et al.CVPR 2023
- REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak SupervisionMihai Fieraru, Mihai Zanfir, Teodor Alexandru Szente, Eduard Gabriel Bazavan et al.NeurIPS 2021 · 44 citations
- U4D: Unsupervised 4D Dynamic Scene UnderstandingArmin Mustafa, Chris Russell, Adrian HiltonICCV 2019 · 6 citations
- Novel View Synthesis of Human Interactions from Sparse Multi-view VideosQing Shuai, Chen Geng, Qi Fang, Sida Peng et al.SIGGRAPH 2022 · 43 citations
- Neural Reconstruction of Relightable Human Model from Monocular VideoWenzhang Sun, Yunlong Che, Yandong Guo, Han HuangICCV 2023 · 20 citations
