Class-agnostic Reconstruction of Dynamic Objects from Videos
Zhongzheng Ren, Xiaoming Zhao, Alexander G. Schwing
Abstract
We introduce REDO, a class-agnostic framework to REconstruct the Dynamic Objects from RGBD or calibrated videos. Compared to prior work, our problem setting is more realistic yet more challenging for three reasons: 1) due to occlusion or camera settings an object of interest may never be entirely visible, but we aim to reconstruct the complete shape; 2) we aim to handle different object dynamics including rigid motion, non-rigid motion, and articulation; 3) we aim to reconstruct different categories of objects with one unified framework. To address these challenges, we develop two novel modules. First, we introduce a canonical 4D implicit function which is pixel-aligned with aggregated temporal visual cues. Second, we develop a 4D transformation module which captures object dynamics to support temporal propagation and aggregation. We study the efficacy of REDO in extensive experiments on synthetic RGBD video datasets SAIL-VOS 3D and DeformingThings4D++, and on real-world video data 3DPW. We find REDO outperforms state-of-the-art dynamic reconstruction methods by a margin. In ablation studies we validate each developed component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoYanqin Jiang, Li Zhang, Jin Gao, Weiming Hu et al.ICLR 2024 · 120 citations
- CASA: Category-agnostic Skeletal Animal ReconstructionYuefan Wu, Zeyuan Chen, Shaowei Liu, Zhongzheng Ren et al.NeurIPS 2022 · 47 citations
- GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-MeshJing Wen, Xiaoming Zhao, Zhongzheng Ren, Alexander G. Schwing et al.CVPR 2024 · 33 citations
- Total-Recon: Deformable Scene Reconstruction for Embodied View SynthesisChonghyuk Song, Gengshan Yang, Kangle Deng, Jun-Yan Zhu et al.ICCV 2023 · 27 citations
- Occupancy Planes for Single-View RGB-D Human ReconstructionXiaoming Zhao, Yuan-Ting Hu, Zhongzheng Ren, Alexander G. SchwingAAAI 2023 · 9 citations
Builds on15
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- PolyGen: An Autoregressive Generative Model of 3D MeshesCharlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. BattagliaICML 2020 · 339 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Occupancy Flow: 4D Reconstruction by Learning Particle DynamicsMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerICCV 2019 · 314 citations
- Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. BlackICCV 2019 · 183 citations
Related papers
- 4D Primitive-Mâché: Glueing Primitives for Persistent 4D Scene ReconstructionKirill Mazur, Marwan Taher, Andrew J. DavisonCVPR 2026 · 1 citation
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object InteractionXianghui Xie, Bowen Wen, Yan Chang, Hesam Rabeti et al.CVPR 2026 · 16 citations
- 4DComplete: Non-Rigid Motion Estimation Beyond the Observable SurfaceYang Li, Hikari Takehara, Takafumi Taketomi, Bo Zheng et al.ICCV 2021 · 160 citations
- MoRe: Motion-aware Feed-forward 4D Reconstruction TransformerJuntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di et al.CVPR 2026 · 8 citations
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosZhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou et al.NeurIPS 2025 · 51 citations
