Root Pose Decomposition Towards Generic Non-rigid 3D Reconstruction with Monocular Videos
Yikai Wang, Yinpeng Dong, Fuchun Sun, Xiao Yang
Abstract
This work focuses on the 3D reconstruction of non-rigid objects based on monocular RGB video sequences. Concretely, we aim at building high-fidelity models for generic object categories and casually captured scenes. To this end, we do not assume known root poses of objects, and do not utilize category-specific templates or dense pose priors. The key idea of our method, Root Pose Decomposition (RPD), is to maintain a per-frame root pose transformation, meanwhile building a dense field with local transformations to rectify the root pose. The optimization of local transformations is performed by point registration to the canonical space. We also adapt RPD to multi-object scenarios with object occlusions and individual differences. As a result, RPD allows non-rigid 3D reconstruction for complicated scenarios containing objects with large deformations, complex motion patterns, occlusions, and scale diversities of different individuals. Such a pipeline potentially scales to diverse sets of objects in the wild. We experimentally show that RPD surpasses state-of-the-art methods on the challenging DAVIS, OVIS, and AMA datasets. We provide video results in https://rpd-share.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian SurfelsYikai Wang, Xinzhou Wang, Zilong Chen, Zhengyi Wang et al.NeurIPS 2024 · 40 citations
- SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian SplattingJun-Jee Chao, Volkan IslerCVPR 2026 · 1 citation
- OSN: Infinite Representations of Dynamic 3D Scenes from Monocular VideosZiyang Song, Jinxi Li, Bo YangICML 2024
Builds on14
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
- Lepard: Learning partial point cloud matching in rigid and deformable scenesYang Li, Tatsuya HaradaCVPR 2022 · 163 citations
- Continuous Surface EmbeddingsNatalia Neverova, David Novotný, Marc Szafraniec, Vasil Khalidov et al.NeurIPS 2020 · 116 citations
Related papers
- Total-Recon: Deformable Scene Reconstruction for Embodied View SynthesisChonghyuk Song, Gengshan Yang, Kangle Deng, Jun-Yan Zhu et al.ICCV 2023 · 27 citations
- Towards Robust and Smooth 3D Multi-Person Pose Estimation from Monocular Videos in the WildSungchan Park, Eunyi You, Inhoe Lee, Joonseok LeeICCV 2023 · 16 citations
- Free-Moving Object Reconstruction and Pose Estimation with Virtual CameraHaixin Shi, Yinlin Hu, Daniel Koguciuk, Juan-Ting Lin et al.AAAI 2025 · 2 citations
- Class-agnostic Reconstruction of Dynamic Objects from VideosZhongzheng Ren, Xiaoming Zhao, Alexander G. SchwingNeurIPS 2021 · 11 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
