MoCaNet: Motion Retargeting In-the-Wild via Canonicalization Networks
Wentao Zhu, Zhuoqian Yang, Ziang Di, Wayne Wu, Yizhou Wang, Chen Change Loy
Abstract
We present a novel framework that brings the 3D motion retargeting task from controlled environments to in-the-wild scenarios. In particular, our method is capable of retargeting body motion from a character in a 2D monocular video to a 3D character without using any motion capture system or 3D reconstruction procedure. It is designed to leverage massive online videos for unsupervised training, requiring neither 3D annotations nor motion-body pairing information. The proposed method is built upon two novel canonicalization operations, structure canonicalization and view canonicalization. Trained with the canonicalization operations and the derived regularizations, our method learns to factorize a skeleton sequence into three independent semantic subspaces, i.e., motion, structure, and view angle. The disentangled representation enables motion retargeting from 2D to 3D with high precision. Our method achieves superior performance on motion transfer benchmarks with large body variations and challenging actions. Notably, the canonicalized skeleton sequence could serve as a disentangled and interpretable representation of human motion that benefits action analysis and motion retrieval 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bc020c1-cfac-46ce-92ac-2530a9dc6b58Cited by top-tier papers4
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsSicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li et al.ACM MM 2023 · 17 citations
- ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis ScoringYuan Xu, Xiaoxuan Ma, Jiajun Su, Wentao Zhu et al.CVPR 2024 · 6 citations
- ReActor: Reinforcement Learning for Physics-Aware Motion RetargetingDavid Müller, Agon Serifi, Sammy Christen, Ruben Grandia et al.SIGGRAPH 2026
Builds on15
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 368 citations
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- Optimizing Network Structure for 3D Human Pose EstimationHai Ci, Chunyu Wang, Xiaoxuan Ma, Yizhou WangICCV 2019 · 267 citations
Related papers
- TransMoMo: Invariance-Driven Unsupervised Video Motion RetargetingZhuoqian Yang, Wentao Zhu, Wayne Wu, Chen Qian et al.CVPR 2020
- TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-AnimationCheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao et al.SIGGRAPH 2026
- Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose EstimationJogendra Nath Kundu, Siddharth Seth, Rahul M. V., Mugalodi Rakesh et al.AAAI 2020 · 57 citations
- Text-to-Any-Skeleton Motion Generation Without RetargetingQingyuan Liu, Ke Lu, Kun Dong, Jian Xue et al.ICCV 2025 · 1 citation
- Motion4Motion: Motion Transfer Across Subjects at InferenceLing-Hao Chen, Zixin Yin, Duomin Wang, Xianfang Zeng et al.SIGGRAPH 2026
