Text2Performer: Text-Driven Human Video Generation
Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu, Chen Change Loy, Ziwei Liu
摘要
Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the appearance and motions of a target performer. Compared to general text-driven video generation, human-centric video generation requires maintaining the appearance of synthesized human while performing complex motions. In this work, we present Text2Performer to generate vivid human videos with articulated motions from texts. Text2Performer has two novel designs: 1) decomposed human representation and 2) diffusion-based motion sampler. First, we decompose the VQVAE latent space into human appearance and pose representation in an unsupervised manner by utilizing the nature of human videos. In this way, the appearance is well maintained along the generated frames. Then, we propose continuous VQ-diffuser to sample a sequence of pose embeddings. Unlike existing VQ-based methods that operate in the discrete space, continuous VQ-diffuser directly outputs the continuous pose embeddings for better motion modeling. Finally, motion-aware masking strategy is designed to mask the pose embeddings spatial-temporally to enhance the temporal coherence. Moreover, to facilitate the task of text-driven human video generation, we contribute a Fashion-Text2Video dataset with manually annotated action labels and text descriptions. Extensive experiments demonstrate that Text2Performer generates high-quality human videos (up to 512 × 256 resolution) with diverse appearances and flexible motions. Our project page is https://yumingj.github.io/projects/Text2Performer.html
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng 等NeurIPS 2024 · 被引用 291 次
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 被引用 113 次
- Disco: Disentangled Control for Realistic Human Dance GenerationTan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai 等CVPR 2024 · 被引用 62 次
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li 等ICLR 2024 · 被引用 34 次
- Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and PracticalityTianle Zhang, Langtian Ma, Yuchen Yan, Yuchen Zhang 等NeurIPS 2024 · 被引用 8 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
相关 Paper
- DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationChenyang Wang, Zerong Zheng, Tao Yu, Xiaoqian Lv 等CVPR 2024 · 被引用 3 次
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 被引用 2 次
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao 等ICML 2026 · 被引用 9 次
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan 等ICCV 2025 · 被引用 11 次
- RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal ReperformanceYuheng Jiang, Zhehao Shen, Chengcheng Guo, Yu Hong 等CVPR 2025
