S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
Hao Zhang, Fang Li, Samyak Rawlekar, Narendra Ahuja
摘要
Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational resources and training time, and require additional human annotations such as predefined parametric models, camera poses, and key points, limiting their generalizability. We propose Synergistic Shape and Skeleton Optimization (S3O), a novel two-phase method that forgoes these prerequisites and efficiently learns parametric models including visible shapes and underlying skeletons. Conventional strategies typically learn all parameters simultaneously, leading to interdependencies where a single incorrect prediction can result in significant errors. In contrast, S3O adopts a phased approach: it first focuses on learning coarse parametric models, then progresses to motion learning and detail addition. This method substantially lowers computational complexity and enhances robustness in reconstruction from limited viewpoints, all without requiring additional annotations. To address the current inadequacies in 3D reconstruction from monocular video benchmarks, we collected the PlanetZoo dataset. Our experimental evaluations on standard benchmarks and the PlanetZoo dataset affirm that S3O provides more accurate 3D reconstruction, and plausible skeletons, and reduces the training time by approximately 60% compared to the state-of-theart, thus advancing the state of the art in dynamic object reconstruction. The code is available on GitHub at: https://github.com/haoz19/LIMR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference OptimizationJingfeng Guo, Jian Liu, Jinnan Chen, Shiwei Mao 等NeurIPS 2025 · 被引用 8 次
- RigMo: Unifying Rig and Motion Learning for Generative AnimationHao Zhang, Jiahao Luo, Bohui Wan, Yizhou Zhao 等CVPR 2026 · 被引用 6 次
- GestureHYDRA: Semantic Co-Speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented GenerationQuanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan 等ICCV 2025 · 被引用 5 次
- ARMO: Autoregressive Rigging for Multi-Category ObjectsMingze Sun, Shiwei Mao, Keyi Chen, Yurun Chen 等ICCV 2025 · 被引用 3 次
- DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive CharactersMingze Sun, Junhao Chen, Junting Dong, Yurun Chen 等CVPR 2025
它引用的顶会 Paper14
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 被引用 789 次
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoChung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron 等CVPR 2022 · 被引用 411 次
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa 等ICCV 2023 · 被引用 390 次
- Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. BlackICCV 2019 · 被引用 183 次
相关 Paper
- Learning Implicit Representation for Reconstructing Articulated ObjectsHao Zhang, Fang Li, Samyak Rawlekar, Narendra AhujaICLR 2024 · 被引用 12 次
- Learning Articulated Shape with Keypoint Pseudo-Labels from Web ImagesAnastasis Stathopoulos, Georgios Pavlakos, Ligong Han, Dimitris N. MetaxasCVPR 2023
- SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian SplattingJun-Jee Chao, Volkan IslerCVPR 2026 · 被引用 1 次
- Towards Unstructured Unlabeled Optical Mocap: A Video Helps!Nicholas Milef, John Keyser, Shu KongSIGGRAPH 2024 · 被引用 2 次
- SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D ObservationCan Zhang, Gim Hee LeeCVPR 2026
