BANMo: Building Animatable 3D Neural Models from Many Casual Videos
Gengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan, Andrea Vedaldi, Hanbyul Joo
Abstract
Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It aims to create high-fidelity, articulated 3D models from many casual RGB videos in a differentiable rendering framework. Our key in-sight is to merge three schools of thought: (1) classic deformable shape models that make use of articulated bones and blend skinning, (2) canonical embeddings that establish correspondences between pixels and a canonical 3D model, and (3) volumetric neural radiance fields (NeRFs) that are amenable to gradient-based optimization. We introduce neural blend skinning models that allow for differentiable and invertible articulated deformations. When combined with canonical embeddings, such models allow us to establish dense correspondences across videos that can be self-supervised with cycle consistency. On real and synthetic datasets, our method shows higher-fidelity 3D reconstructions than prior works for humans and animals, with the ability to render realistic images from novel viewpoints. Project page: https://banmo-www.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0558aeaf-c903-4ffb-9310-97ea98e1cae1Cited by top-tier papers30
- MegaPortraits: One-shot Megapixel Neural Head AvatarsNikita Drobyshev, Jenya Chelishev, Taras Khakhulin, Aleksei Ivakhnenko et al.ACM MM 2022 · 86 citations
- DynPoint: Dynamic Neural Point For View SynthesisKaichen Zhou, Jia-Xing Zhong, Sangyun Shin, Kai Lu et al.NeurIPS 2023 · 46 citations
- Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian SurfelsYikai Wang, Xinzhou Wang, Zilong Chen, Zhengyi Wang et al.NeurIPS 2024 · 40 citations
- Watch It Move: Unsupervised Discovery of 3D Joints for Re-Posing of Articulated ObjectsAtsuhiro Noguchi, Umar Iqbal, Jonathan Tremblay, Tatsuya Harada et al.CVPR 2022 · 32 citations
- Artemis: articulated neural pets with appearance and motion synthesisHaimin Luo, Teng Xu, Yuheng Jiang, Chenglin Zhou et al.SIGGRAPH 2022 · 30 citations
Builds on30
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
Related papers
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 316 citations
- Animatable Neural Radiance Fields for Modeling Dynamic Human BodiesSida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang et al.ICCV 2021 · 461 citations
- Template-free Articulated Neural Point Clouds for Reposable View SynthesisLukas Uzolas, Elmar Eisemann, Petr KellnhoferNeurIPS 2023 · 20 citations
- Learning Compositional Radiance Fields of Dynamic Human HeadsZiyan Wang, Timur M. Bagautdinov, Stephen Lombardi, Tomas Simon et al.CVPR 2021
- REACTO: Reconstructing Articulated Objects from a Single VideoChaoyue Song, Jiacheng Wei, Chuan Sheng Foo, Guosheng Lin et al.CVPR 2024 · 7 citations
