SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
Shuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang, Zhuoxin Liu, Ke Ma, Yan Wang, Kecheng Zheng, Xing Zhu, Yujun Shen, Hengshuang Zhao
Abstract
Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arbitrary subjects transfer through precise textual conditioning. Existing approaches often rely on semantic-level alignment, expecting the model to learn new motion concepts and combine them with other entities (e.g.,''cats''or''dogs'') to produce visually appealing results. However, video data involve complex spatio-temporal patterns, and focusing solely on semantics cause the model to overlook the visual complexity of motion. Conversely, tuning only the visual representation leads to semantic confusion in representing the intended action. To address these limitations, we propose SynMotion, a new motion-customized video generation model that jointly leverages semantic guidance and visual adaptation. At the semantic level, we introduce the dual-embedding semantic comprehension mechanism which disentangles subject and motion representations, allowing the model to learn customized motion features while preserving its generative capabilities for diverse subjects. At the visual level, we integrate parameter-efficient motion adapters into a pre-trained video generation model to enhance motion fidelity and temporal coherence. Furthermore, we introduce a new embedding-specific training strategy which alternately optimizes subject and motion embeddings, supported by the manually constructed Subject Prior Video (SPV) training dataset. This strategy promotes motion specificity while preserving generalization across diverse subjects. Lastly, we introduce MotionBench, a newly curated benchmark with diverse motion patterns. Experimental results across both T2V and I2V settings demonstrate that outperforms existing baselines. Project page: https://lucaria-academy.github.io/SynMotion/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f6a9784-03be-43a0-ade2-31fd59a07227Cited by top-tier papers2
- FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesShuai Tan, Bill Gong, Bin Ji, Ye PanICCV 2025 · 3 citations
- Stylized-Face: A Million-Level Stylized Face Dataset for Face RecognitionZhengyuan Peng, Jianqing Xu, Yuge Huang, Jinkun Hao et al.ICCV 2025 · 1 citation
Builds on58
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Versatile Transition Generation with Image-to-Video DiffusionZuhao Yang, Jiahui Zhang, Yingchen Yu, Shijian Lu et al.ICCV 2025
- Space-Time Diffusion Features for Zero-Shot Text-Driven Motion TransferDanah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten et al.CVPR 2024 · 29 citations
- DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video CustomizationWenchuan Wang, Mengqi Huang, Yijing Tu, Zhendong MaoICCV 2025 · 3 citations
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion ModelsAritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amir Habibian et al.ICLR 2026 · 15 citations
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang et al.ICCV 2025 · 1 citation
