Dynamic Concepts Personalization from Single Videos
Rameen Abdal, Or Patashnik, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Sergey Tulyakov, Daniel Cohen-Or, Kfir Aberman
Abstract
Fig. 1. We personalize a video model to capture dynamic concepts -entities defined not only by their appearance but also by their unique motion patterns, such as the fluid motion of ocean waves or the flickering dynamics of a bonfire (left). This enables high-fidelity generation, editing, and the composition of these dynamic elements into a single video, where they interact naturally (right).
Personalizing generative text-to-image models has seen remarkable progress, but extending this personalization to text-to-video models presents unique challenges. Unlike static concepts, personalizing text-to-video models has the potential to capture dynamic concepts -entities defined not only by their appearance but also by their motion. In this paper, we introduce Setand-Sequence, a novel framework for personalizing Diffusion Transformers (DiTs)-based generative video models with dynamic concepts. Our approach imposes a spatio-temporal weight space within an architecture that does not explicitly separate spatial and temporal features. This is achieved in two key stages. First, we fine-tune Low-Rank Adaptation (LoRA) layers using an unordered set of frames from the video to learn an identity LoRA basis that represents the appearance, free from temporal interference. In the second stage, with the identity LoRAs frozen, we augment their coefficients with Motion Residuals and fine-tune them on the full video sequence, capturing motion dynamics. Our Set-and-Sequence framework resulting in a spatio-temporal weight space effectively embeds dynamic concepts into the video model's output domain, enabling unprecedented editability and compositionality, and setting a new benchmark for personalizing dynamic concepts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled FinetuningYue Ma, Yulong Liu, Qiyuan Zhu, Xiangpeng Yang et al.ICLR 2026 · 70 citations
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image GenerationYuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa et al.CVPR 2026 · 10 citations
- Composing Concepts from Images and Videos via Concept-prompt BindingXianghao Kong, Zeyu Zhang, Yuwei Guo, Zhuoran Zhao et al.CVPR 2026 · 2 citations
- FlowMotion: Training-Free Flow Guidance for Video Motion TransferZhen Wang, Youcan Xu, Jun Xiao, Long ChenCVPR 2026 · 1 citation
- Visual Personalization Turing TestRameen Abdal, James Burgess, Sergey Tulyakov, Kuan-Chieh Jackson WangCVPR 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- CustomTTT: Motion and Appearance Customized Video Generation via Test-Time TrainingXiuli Bi, Jian Lu, Bo Liu, Xiaodong Cun et al.AAAI 2025 · 10 citations
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video CreationWenhui Song, Hanhui Li, Jiehui Huang, Panwen Hu et al.ACM MM 2025
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition AbilitiesTao Wu, Yong Zhang, Xintao Wang, Xianpan Zhou et al.AAAI 2025 · 62 citations
