Flow Guided Transformable Bottleneck Networks for Motion Retargeting
Jian Ren, Menglei Chai, Oliver J. Woodford, Kyle Olszewski, Sergey Tulyakov
Abstract
Human motion retargeting aims to transfer the motion of one person in a "driving" video or set of images to another person. Existing efforts leverage a long training video from each target person to train a subject-specific motion transfer model. However, the scalability of such methods is limited, as each model can only generate videos for the given target subject, and such training videos are labor-intensive to acquire and process. Few-shot motion transfer techniques, which only require one or a few images from a target, have recently drawn considerable attention. Methods addressing this task generally use either 2D or explicit 3D representations to transfer motion, and in doing so, sacrifice either accurate geometric modeling or the flexibility of an end-to-end learned representation. Inspired by the Transformable Bottleneck Network, which renders novel views and manipulations of rigid objects, we propose an approach based on an implicit volumetric representation of the image content, which can then be spatially manipulated using volumetric flow fields. We address the challenging question of how to aggregate information across different body poses, learning flow fields that allow for combining content from the appropriate regions of input images of highly non-rigid human subjects performing complex motions into a single implicit volumetric representation. This allows us to learn our 3D representation solely from videos of moving people. Armed with both 3D object understanding and end-to-end learned rendering, this categorically novel representation delivers state-of-the-art image generation quality, as shown by our quantitative and qualitative evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca15d6bf-ea5c-4726-97d3-96d4fb398e14Cited by top-tier papers10
- AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose ControlRuixiang Jiang, Can Wang, Jingbo Zhang, Menglei Chai et al.ICCV 2023 · 103 citations
- Show Me What and Tell Me How: Video Synthesis via Multimodal ConditioningLigong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri et al.CVPR 2022 · 36 citations
- Structure-Aware Motion Transfer with Deformable Anchor ModelJiale Tao, Biao Wang, Borun Xu, Tiezheng Ge et al.CVPR 2022 · 33 citations
- Adaptive Affine Transformation: A Simple and Effective Operation for Spatial Misaligned Image GenerationZhimeng Zhang, Yu DingACM MM 2022 · 18 citations
- BodyGAN: General-purpose Controllable Neural Human Body GenerationChaojie Yang, Hanhui Li, Shengjie Wu, Shengkai Zhang et al.CVPR 2022 · 8 citations
Builds on19
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
Related papers
- Transformable Bottleneck NetworksKyle Olszewski, Sergey Tulyakov, Oliver J. Woodford, Hao Li et al.ICCV 2019 · 79 citations
- Few-Shot Human Motion Transfer by Personalized Geometry and Texture ModelingZhichao Huang, Xintong Han, Jia Xu, Tong ZhangCVPR 2021
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoChung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron et al.CVPR 2022 · 411 citations
- TransMoMo: Invariance-Driven Unsupervised Video Motion RetargetingZhuoqian Yang, Wentao Zhu, Wayne Wu, Chen Qian et al.CVPR 2020
- MoCaNet: Motion Retargeting In-the-Wild via Canonicalization NetworksWentao Zhu, Zhuoqian Yang, Ziang Di, Wayne Wu et al.AAAI 2022 · 24 citations
