ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Ce Liu, Deva Ramanan
摘要
We introduce ViSER, a method for recovering articulated 3D shapes and dense 3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple synchronized cameras, strong category-specific priors, or 2D keypoint supervision. We show that none of these are required if one can reliably estimate long-range correspondences in a video, making use of only 2D object masks and two-frame optical flow as inputs. ViSER infers correspondences by matching 2D pixels to a canonical, deformable 3D mesh via video-specific surface embeddings that capture the view-independent appearance features of each surface point. These embeddings behave as a continuous set of keypoint descriptors defined over the mesh surface, which can be used to establish dense long-range correspondences across pixels. The surface embeddings are implemented as coordinate-based MLPs that are fit to each video via self-supervised losses. Experimental results show that ViSER compares favorably against prior work on challenging videos of humans with loose clothing and unusual poses as well as animal videos from DAVIS and YTVOS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li 等ICCV 2023 · 被引用 238 次
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan 等CVPR 2022 · 被引用 113 次
- LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part DiscoveryChun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein 等NeurIPS 2022 · 被引用 83 次
- Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D CameraHongrui Cai, Wanquan Feng, Xuetao Feng, Yan Wang 等NeurIPS 2022 · 被引用 83 次
- CASA: Category-agnostic Skeletal Animal ReconstructionYuefan Wu, Zeyuan Chen, Shaowei Liu, Zhongzheng Ren 等NeurIPS 2022 · 被引用 47 次
它引用的顶会 Paper9
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- Continuous Surface EmbeddingsNatalia Neverova, David Novotný, Marc Szafraniec, Vasil Khalidov 等NeurIPS 2020 · 被引用 116 次
- Deep Non-Rigid Structure From MotionChen Kong, Simon LuceyICCV 2019 · 被引用 72 次
- Online Adaptation for Consistent Mesh Reconstruction in the WildXueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim 等NeurIPS 2020 · 被引用 62 次
- VIBE: Video Inference for Human Body Pose and Shape EstimationMuhammed Kocabas, Nikos Athanasiou, Michael J. BlackCVPR 2020
相关 Paper
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani 等CVPR 2022 · 被引用 111 次
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 被引用 1 次
- Vid2Avatar: 3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene DecompositionChen Guo, Tianjian Jiang, Xu Chen, Jie Song 等CVPR 2023
- Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstructionDavid Novotný, Roman Shapovalov, Andrea VedaldiNeurIPS 2020 · 被引用 9 次
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular VideoBoyi Jiang, Yang Hong, Hujun Bao, Juyong ZhangCVPR 2022 · 被引用 142 次
