Online Adaptation for Consistent Mesh Reconstruction in the Wild
Xueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim, Xiaolong Wang, Ming-Hsuan Yang, Jan Kautz
Abstract
This paper presents an algorithm to reconstruct temporally consistent 3D meshes of deformable object instances from videos in the wild. Without requiring annotations of 3D mesh, 2D keypoints, or camera pose for each video frame, we pose videobased reconstruction as a self-supervised online adaptation problem applied to any incoming test video. We first learn a category-specific 3D reconstruction model from a collection of single-view images of the same category that jointly predicts the shape, texture, and camera pose of an image. Then, at inference time, we adapt the model to a test video over time using self-supervised regularization terms that exploit temporal consistency of an object instance to enforce that all reconstructed meshes share a common texture map, a base shape, as well as parts. We demonstrate that our algorithm recovers temporally consistent and reliable 3D structures from videos of non-rigid objects including those of animals captured in the wild -an extremely challenging task rarely addressed before. Codes and other resources will be maintained at https://sites.google.com/nvidia.com/vmr-2020 . Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dabf702f-d6c9-4b3b-aaa2-2d34915faf1fCited by top-tier papers31
- NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the WildJason Y. Zhang, Gengshan Yang, Shubham Tulsiani, Deva RamananNeurIPS 2021 · 180 citations
- A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape RepresentationJiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille et al.ICCV 2021 · 138 citations
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan et al.CVPR 2022 · 113 citations
- ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape ReconstructionGengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic et al.NeurIPS 2021 · 103 citations
- Test-Time Personalization with a Transformer for Human Pose EstimationYizhuo Li, Miao Hao, Zonglin Di, Nitesh B. Gundavarapu et al.NeurIPS 2021 · 58 citations
Builds on12
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Pixel2Mesh++: Multi-View 3D Mesh Generation via DeformationChao Wen, Yinda Zhang, Zhuwen Li, Yanwei FuICCV 2019 · 279 citations
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang et al.ICCV 2019 · 218 citations
- Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. BlackICCV 2019 · 183 citations
Related papers
- Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the WildKaifeng Zhang, Yang Fu, Shubhankar Borse, Hong Cai et al.ICLR 2023 · 8 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Learning Monocular 3D Reconstruction of Articulated Categories From MotionFilippos Kokkinos, Iasonas KokkinosCVPR 2021
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu et al.CVPR 2020
- Self-Supervised 3D Mesh Reconstruction From Single ImagesTao Hu, Liwei Wang, Xiaogang Xu, Shu Liu et al.CVPR 2021
