Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free Videos
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, Qifeng Chen
Abstract
Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose captions and the generative prior models for videos. In this work, we design a novel two-stage training scheme that can utilize easily obtained datasets (i.e., image pose pair and pose-free video) and the pre-trained text-to-image (T2I) model to obtain the pose-controllable character videos. Specifically, in the first stage, only the keypoint image pairs are used only for a controllable text-to-image generation. We learn a zero-initialized convolutional encoder to encode the pose information. In the second stage, we finetune the motion of the above network via a pose-free video dataset by adding the learnable temporal self-attention and reformed cross-frame self-attention blocks. Powered by our new designs, our method successfully generates continuously pose-controllable character videos while keeps the editing and concept composition ability of the pre-trained T2I model. The code and models are available on https://follow-your-pose.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6060fa11-086c-42e1-ba55-ebbe6b66fcfdCited by top-tier papers144
- ControlVideo: Training-free Controllable Text-to-video GenerationYabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang et al.ICLR 2024 · 359 citations
- Preserve Your Own Correlation: A Noise Prior for Video Diffusion ModelsSongwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon et al.ICCV 2023 · 319 citations
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen et al.ICLR 2024 · 175 citations
- PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeXuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu et al.ICLR 2024 · 166 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingJianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo et al.ACM MM 2025 · 6 citations
- Video-P2P: Video Editing with Cross-Attention ControlShaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin et al.CVPR 2024 · 99 citations
- DPE: Disentanglement of Pose and Expression for General Video Portrait EditingYouxin Pang, Yong Zhang, Weize Quan, Yanbo Fan et al.CVPR 2023
- Text2Performer: Text-Driven Human Video GenerationYuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu et al.ICCV 2023 · 74 citations
- IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot MannerYuyang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang et al.CVPR 2025
