CASA: Category-agnostic Skeletal Animal Reconstruction
Yuefan Wu, Zeyuan Chen, Shaowei Liu, Zhongzheng Ren, Shenlong Wang
Abstract
Recovering the skeletal shape of an animal from a monocular video is a longstanding challenge. Prevailing animal reconstruction methods often adopt a control-point driven animation model and optimize bone transforms individually without considering skeletal topology, yielding unsatisfactory shape and articulation. In contrast, humans can easily infer the articulation structure of an unknown animal by associating it with a seen articulated character in their memory. Inspired by this fact, we present CASA, a novel Category-Agnostic Skeletal Animal reconstruction method consisting of two major components: a video-to-shape retrieval process and a neural inverse graphics framework. During inference, CASA first retrieves an articulated shape from a 3D character assets bank so that the input video scores highly with the rendered image, according to a pretrained language-vision model. CASA then integrates the retrieved character into an inverse graphics framework and jointly infers the shape deformation, skeleton structure, and skinning weights through optimization. Experiments validate the efficacy of CASA regarding shape reconstruction and articulation. We further demonstrate that the resulting skeletal-animated characters can be used for re-animation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- Template-free Articulated Neural Point Clouds for Reposable View SynthesisLukas Uzolas, Elmar Eisemann, Petr KellnhoferNeurIPS 2023 · 20 citations
- Template-free Articulated Gaussian Splatting for Real-time Reposable Dynamic View SynthesisDiwen Wan, Yuxiang Wang, Ruijie Lu, Gang ZengNeurIPS 2024 · 18 citations
- Learning Implicit Representation for Reconstructing Articulated ObjectsHao Zhang, Fang Li, Samyak Rawlekar, Narendra AhujaICLR 2024 · 12 citations
- Bayesian Diffusion Models for 3D Shape ReconstructionHaiyang Xu, Yu Lei, Zeyuan Chen, Xiang Zhang et al.CVPR 2024 · 12 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
Related papers
- Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery from Sparse Image EnsembleChun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein et al.CVPR 2023
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular VideosKehong Gong, Zhengyu Wen, Xiaoyu He, Mingxi Xu et al.CVPR 2026 · 8 citations
- Distilling Neural Fields for Real-Time Articulated Shape ReconstructionJeff Tan, Gengshan Yang, Deva RamananCVPR 2023
- Semantic-Aware Motion Encoding for Topology-Agnostic Character AnimationZongye Zhang, Yuzhuo Cui, Qingjie Liu, Yunhong WangICML 2026 · 1 citation
- Reconstructing Animatable Categories from VideosGengshan Yang, Chaoyang Wang, N. Dinesh Reddy, Deva RamananCVPR 2023
