MotionBooth: Motion-Aware Customized Text-to-Video Generation
Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou, Yining Li, Yunhai Tong, Kai Chen
Abstract
In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a text-to-video model to capture the object's shape and attributes accurately. Our approach presents subject region loss and video preservation loss to enhance the subject's learning performance, along with a subject token cross-attention loss to integrate the customized subject with motion control signals. Additionally, we propose training-free techniques for managing subject and camera motions during inference. In particular, we utilize cross-attention map manipulation to govern subject motion and introduce a novel latent shift module for camera movement control as well. MotionBooth excels in preserving the appearance of subjects while simultaneously controlling the motions in generated videos. Extensive quantitative and qualitative evaluations demonstrate the superiority and effectiveness of our method. Our project page is at https://jianzongwu.github.io/projects/motionbooth
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1595108-1116-46b0-8b4d-a2d810db46a6Cited by top-tier papers52
- MotionStream: Real-Time Video Generation with Interactive Motion ControlsJoonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu et al.ICLR 2026 · 79 citations
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceRuihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang et al.NeurIPS 2025 · 50 citations
- MultiShotMaster: A Controllable Multi-Shot Video Generation FrameworkQinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian et al.CVPR 2026 · 33 citations
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng et al.ICLR 2026 · 32 citations
- Stand-In: A Lightweight and Plug-and-Play Identity Control for Video GenerationBowen Xue, Zheng-Peng Duan, Qixin Yan, Wenjing Wang et al.CVPR 2026 · 28 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- MotionFlow: Attention-Driven Motion Transfer in Video Diffusion ModelsTuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, Pinar YanardagAAAI 2026
- AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image GenerationLianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu et al.NeurIPS 2024 · 18 citations
- ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion TransferJiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng et al.CVPR 2025
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang et al.AAAI 2025 · 1 citation
- Storybooth: Training-Free Multi-Subject Consistency for Improved Visual StorytellingJaskirat Singh, Junshen K. Chen, Jonas Kohler, Michael F. CohenICLR 2025
