TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting
Zhuoqian Yang, Wentao Zhu, Wayne Wu, Chen Qian, Qiang Zhou, Bolei Zhou, Chen Change Loy
Abstract
Abstract We present a lightweight video motion retargeting approach TransMoMo that is capable of transferring motion of a person in a source video realistically to another video of a target person (Fig. 1 ). Without using any paired data for supervision, the proposed method can be trained in an unsupervised manner by exploiting invariance properties of three orthogonal factors of variation including motion, structure, and view-angle. Specifically, with loss functions carefully derived based on invariance, we train an autoencoder to disentangle the latent representations of such factors given the source and target video clips. This allows us to selectively transfer motion extracted from the source video seamlessly to the target video in spite of structural and view-angle disparities between the source and the target. The relaxed assumption of paired data allows our method to be trained on a vast amount of videos needless of manual annotation of source-target pairing, leading to improved robustness against large structural variations and extreme motion in videos. We demonstrate the effectiveness of our method over the state-of-the-art methods such as NKN [41] , EDN [8] and LCM [4] . Code, model and data are publicly available on our project page. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54a2e933-c0e5-4ef5-8d81-4a4b4719f511Cited by top-tier papers20
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 219 citations
- Zero-shot Synthesis with Group-Supervised LearningYunhao Ge, Sami Abu-El-Haija, Gan Xin, Laurent IttiICLR 2021 · 45 citations
- FashionMirror: Co-attention Feature-remapping Virtual Try-on with Sequential Template PosesChieh-Yun Chen, Ling Lo, Pin-Jui Huang, Hong-Han Shuai et al.ICCV 2021 · 34 citations
- Structure-Aware Motion Transfer with Deformable Anchor ModelJiale Tao, Biao Wang, Borun Xu, Tiezheng Ge et al.CVPR 2022 · 33 citations
Builds on3
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- Make a Face: Towards Arbitrary High Fidelity Face ManipulationShengju Qian, Kwan-Yee Lin, Wayne Wu, Yangxiaokang Liu et al.ICCV 2019 · 75 citations
Related papers
- MoCaNet: Motion Retargeting In-the-Wild via Canonicalization NetworksWentao Zhu, Zhuoqian Yang, Ziang Di, Wayne Wu et al.AAAI 2022 · 24 citations
- Flow Guided Transformable Bottleneck Networks for Motion RetargetingJian Ren, Menglei Chai, Oliver J. Woodford, Kyle Olszewski et al.CVPR 2021
- MotionShot: Adaptive Motion Transfer Across Arbitrary Objects for Text-to-Video GenerationYanchen Liu, Yanan Sun, Zhening Xing, Junyao Gao et al.ICCV 2025 · 5 citations
- Unpaired motion style transfer from video to animationKfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or et al.SIGGRAPH 2020 · 178 citations
- FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression FeaturesAndre Rochow, Max Schwarz, Sven BehnkeCVPR 2024 · 17 citations
