VMC: Video Motion Customization Using Temporal Attention Adaption for Text-to-Video Diffusion Models
Hyeonho Jeong, Geon Yeong Park, Jong Chul Ye
Abstract
Figure 1. Using only a single video portraying any type of motion, our Video Motion Customization framework allows for generating a wide variety of videos characterized by the same motion but in entirely distinct contexts and better spatial/temporal resolution. 8-frame input videos are translated to 29-frame videos in different contexts while closely following the target motion. The visualized frames for the first video are at indexes 1, 9, and 17. A comprehensive view of these motions in the form of videos can be explored at our project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4aae27df-82d8-4f7b-a6d3-9388819bf8afCited by top-tier papers48
- EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled FinetuningYue Ma, Yulong Liu, Qiyuan Zhu, Xiangpeng Yang et al.ICLR 2026 · 70 citations
- Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion ModelsHyeonho Jeong, Jong Chul YeICLR 2024 · 68 citations
- Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object MotionShiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 46 citations
- FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image EditingJeongsol Kim, Yeobin Hong, Jonghyun Park, Jong Chul YeICLR 2026 · 35 citations
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng et al.ICLR 2026 · 32 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Motion Inversion for Video CustomizationLuozhou Wang, Ziyang Mai, Guibao Shen, Yixun Liang et al.SIGGRAPH 2025 · 2 citations
- VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsChi-Pin Huang, Yen-Siang Wu, Hung-Kai Chung, Kai-Po Chang et al.CVPR 2025
- Spectral Motion Alignment for Video Motion Transfer Using Diffusion ModelsGeon Yeong Park, Hyeonho Jeong, Sang Wan Lee, Jong Chul YeAAAI 2025 · 21 citations
- Framer: Interactive Frame InterpolationWen Wang, Qiuyu Wang, Kecheng Zheng, Hao Ouyang et al.ICLR 2025
- MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose GuidanceYuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang et al.ICML 2025
