Semantics-Aware Motion Retargeting with Vision-Language Models
Haodong Zhang, Zhike Chen, Haocheng Xu, Lei Hao, Xiaofei Wu, Songcen Xu, Zhensong Zhang, Yue Wang, Rong Xiong
Abstract
Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However, most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here, we present a novel Semantics-aware Motion reTargeting (SMT) method with the advantage of vision-language models to extract and maintain meaningful motion semantics. We utilize a differentiable module to ren-der 3D motions. Then the high-level motion semantics are incorporated into the motion retargeting process by feeding the vision-language model with the rendered images and aligning the extracted semantic embeddings. To en-sure the preservation of fine-grained motion details and high-level semantics, we adopt a two-stage pipeline consisting of skeleton-aware pretraining and fine-tuning with semantics and geometry constraints. Experimental results show the effectiveness of the proposed method in producing high-quality motion retargeting results while accurately preserving motion semantics. Project page can be found at https://sites.google.com/view/smtnet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Skinned Motion Retargeting with Dense Geometric Interaction PerceptionZijie Ye, Jia-Wei Liu, Jia Jia, Shikun Sun et al.NeurIPS 2024 · 22 citations
- STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency ConstraintsXiaohang Yang, Qing Wang, Jiahao Yang, Gregory G. Slabaugh et al.ICCV 2025 · 2 citations
- Text-to-Any-Skeleton Motion Generation Without RetargetingQingyuan Liu, Ke Lu, Kun Dong, Jian Xue et al.ICCV 2025 · 1 citation
- AniMimic: Imitating 3D Animation from Video PriorsTianyi Xie, Yunuo Chen, Yaowei Guo, Yin Yang et al.CVPR 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
Related papers
- Skinned Motion Retargeting with Spatially Adaptive Interaction GuidanceSoojin Choi, Seokhyeon Hong, Chaelin Kim, Junghyun Nam et al.SIGGRAPH 2026
- Skinned Motion Retargeting with Residual Perception of Motion Semantics & GeometryJiaxu Zhang, Junwu Weng, Di Kang, Fang Zhao et al.CVPR 2023
- Semantic-Aware Motion Encoding for Topology-Agnostic Character AnimationZongye Zhang, Yuzhuo Cui, Qingjie Liu, Yunhong WangICML 2026 · 1 citation
- Motion-Aligned Word Embeddings for Text-to-Motion GenerationKe Han, Yueming Lyu, Nicu SebeICLR 2026
- Contact-Aware Retargeting of Skinned MotionRuben Villegas, Duygu Ceylan, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 52 citations
