MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
Kehong Gong, Zhengyu Wen, Xiaoyu He, Mingxi Xu, Qi WANG, ning Zhang, Zhengyu Li, Dongze Lian, Wei Zhao, He Xiaoyu, Mingyuan Zhang
摘要
Motion capture now underpins content creation far beyond digital humans, yet most pipelines remain species- or template-specific. We formalize this gap as Category-Agnostic Motion Capture (CAMoCap): given a monocular video and an arbitrary rigged 3D asset as a prompt, the goal is to reconstruct a rotation-based animation (e.g., BVH) that directly drives the specific asset. We present MoCapAnything, a reference-guided, factorized framework that first predicts 3D joint trajectories and then recovers asset-specific rotations via constraint-aware Inverse Kinematics (IK) Fitting. MoCapAnything comprises three learnable modules and a lightweight IK stage: a Reference Prompt Encoder that distills per-joint queries from the asset’s skeleton, mesh, and rendered image set; a Video Feature Extractor that computes dense visual descriptors and reconstructs a coarse 4D deforming mesh to bridge the modality gap between RGB tokens and the point-cloud–like joint space; and a Unified Motion Decoder that fuses these cues to produce temporally coherent trajectories. We also curate Truebones Zoo with 1,038 motion clips, each providing a standardized skeleton–mesh–rendered-video triad. Experiments on in-domain benchmarks and in-the-wild videos show that delivers high-quality skeletal animations and exhibits non-trivial cross-species retargeting across heterogeneous rigs, offering a scalable path toward prompt-based 3D motion capture for arbitrary assets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow: R-DMeshZijie Wu, Lixin Xu, Puhua Jiang, Sicong Liu 等SIGGRAPH 2026
- TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-AnimationCheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao 等SIGGRAPH 2026
它引用的顶会 Paper26
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa 等ICCV 2023 · 被引用 390 次
- Human Pose Regression with Residual Log-likelihood EstimationJiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang 等ICCV 2021 · 被引用 286 次
- End-to-End Multi-Person Pose Estimation with TransformersDahu Shi, Xing Wei, Liangqi Li, Ye Ren 等CVPR 2022 · 被引用 147 次
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani 等CVPR 2022 · 被引用 111 次
相关 Paper
- Semantic-Aware Motion Encoding for Topology-Agnostic Character AnimationZongye Zhang, Yuzhuo Cui, Qingjie Liu, Yunhong WangICML 2026 · 被引用 1 次
- RigAnything: Template-Free Autoregressive Rigging for Diverse 3D AssetsIsabella Liu, Zhan Xu, Wang Yifan, Hao Tan 等SIGGRAPH 2025 · 被引用 11 次
- RigMo: Unifying Rig and Motion Learning for Generative AnimationHao Zhang, Jiahao Luo, Bohui Wan, Yizhou Zhao 等CVPR 2026 · 被引用 6 次
- How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary ObjectsWonkwang Lee, Jongwon Jeong, Taehong Moon, Hyeon-Jong Kim 等ICML 2025
- Precise Action-to-Video Generation Through Visual Action PromptsYuang Wang, Chao Wen, Haoyu Guo, Sida Peng 等ICCV 2025 · 被引用 1 次
