ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Ce Liu, Deva Ramanan
Abstract
We introduce ViSER, a method for recovering articulated 3D shapes and dense 3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple synchronized cameras, strong category-specific priors, or 2D keypoint supervision. We show that none of these are required if one can reliably estimate long-range correspondences in a video, making use of only 2D object masks and two-frame optical flow as inputs. ViSER infers correspondences by matching 2D pixels to a canonical, deformable 3D mesh via video-specific surface embeddings that capture the view-independent appearance features of each surface point. These embeddings behave as a continuous set of keypoint descriptors defined over the mesh surface, which can be used to establish dense long-range correspondences across pixels. The surface embeddings are implemented as coordinate-based MLPs that are fit to each video via self-supervised losses. Experimental results show that ViSER compares favorably against prior work on challenging videos of humans with loose clothing and unusual poses as well as animal videos from DAVIS and YTVOS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb3435cf-1b0a-4d67-acfb-7a131638de77Cited by top-tier papers43
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan et al.CVPR 2022 · 113 citations
- LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part DiscoveryChun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein et al.NeurIPS 2022 · 83 citations
- Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D CameraHongrui Cai, Wanquan Feng, Xuetao Feng, Yan Wang et al.NeurIPS 2022 · 83 citations
- CASA: Category-agnostic Skeletal Animal ReconstructionYuefan Wu, Zeyuan Chen, Shaowei Liu, Zhongzheng Ren et al.NeurIPS 2022 · 47 citations
Builds on9
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Continuous Surface EmbeddingsNatalia Neverova, David Novotný, Marc Szafraniec, Vasil Khalidov et al.NeurIPS 2020 · 116 citations
- Deep Non-Rigid Structure From MotionChen Kong, Simon LuceyICCV 2019 · 72 citations
- Online Adaptation for Consistent Mesh Reconstruction in the WildXueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim et al.NeurIPS 2020 · 62 citations
- VIBE: Video Inference for Human Body Pose and Shape EstimationMuhammed Kocabas, Nikos Athanasiou, Michael J. BlackCVPR 2020
Related papers
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
- Vid2Avatar: 3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene DecompositionChen Guo, Tianjian Jiang, Xu Chen, Jie Song et al.CVPR 2023
- Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstructionDavid Novotný, Roman Shapovalov, Andrea VedaldiNeurIPS 2020 · 9 citations
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 142 citations
