Human MotionFormer: Transferring Human Motions with Vision Transformers
Hongyu Liu, Xintong Han, Chenbin Jin, Lihui Qian, Huawei Wei, Zhe Lin, Faqiang Wang, Haoye Dong, Yibing Song, Jia Xu, Qifeng Chen
摘要
Human motion transfer aims to transfer motions from a target dynamic person to a source static one for motion synthesis. An accurate matching between the source person and the target motion in both large and subtle motion changes is vital for improving the transferred motion quality. In this paper, we propose Human MotionFormer, a hierarchical ViT framework that leverages global and local perceptions to capture large and subtle motion matching, respectively. It consists of two ViT encoders to extract input features (i.e., a target motion image and a source human image) and a ViT decoder with several cascaded blocks for feature matching and motion transfer. In each block, we set the target motion feature as Query and the source person as Key and Value, calculating the crossattention maps to conduct a global feature matching. Further, we introduce a convolutional layer to improve the local perception after the global cross-attention computations. This matching process is implemented in both warping and generation branches to guide the motion transfer. During training, we propose a mutual learning loss to enable the co-supervision between warping and generation branches for better motion representations. Experiments show that our Human MotionFormer sets the new state-of-the-art performance both qualitatively and quantitatively. Project page: https://github.com/KumapowerLIU/ Human-MotionFormer * X. Han and H. Liu contribute equally. † Y. Song and Q. Chen are the corresponding authors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsZhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li 等AAAI 2025 · 被引用 197 次
- Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning MambaHaoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente Carrasco 等NeurIPS 2024 · 被引用 73 次
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu 等ICLR 2026 · 被引用 47 次
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng 等ICLR 2026 · 被引用 32 次
- MotionEditor: Editing Video Motion via Content-Aware DiffusionShuyuan Tu, Qi Dai, Zhi-Qi Cheng, Han Hu 等CVPR 2024 · 被引用 21 次
它引用的顶会 Paper33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
相关 Paper
- Future Motion Dynamic Modeling via Hybrid Supervision for Multi-Person Motion Prediction Uncertainty ReductionYan Zhuang, Yanlu Cai, Weizhong Zhang, Cheng JinACM MM 2024 · 被引用 3 次
- MotionHiFlow: Text-to-Motion via Hierarchical Flow MatchingHeng Li, Xiaotong Lin, Ling-An Zeng, Yulei Kang 等CVPR 2026 · 被引用 7 次
- C2F-FWN: Coarse-to-Fine Flow Warping Network for Spatial-Temporal Consistent Motion TransferDongxu Wei, Xiaowei Xu, Haibin Shen, Kejie HuangAAAI 2021 · 被引用 23 次
- H-ViT: A Hierarchical Vision Transformer for Deformable Image RegistrationMorteza Ghahremani, Mohammad Khateri, Bailiang Jian, Benedikt Wiestler 等CVPR 2024
- A Unified Masked Autoencoder with Patchified Skeletons for Motion SynthesisEsteve Valls Mascaro, Hyemin Ahn, Dongheui LeeAAAI 2024 · 被引用 11 次
