Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLM
Yatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang, Shoufa Chen, Chongjian Ge, Peize Sun, Weifeng Chen, Wenqi Shao, Xuefeng Xiao, Weilin Huang, Ping Luo
摘要
Text-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often entails multiple revisions and iterative inference to refine user-provided prompts. Current automatic methods for refining prompts encounter challenges such as Modality-Inconsistency, Cost-Discrepancy, and Model-Unaware when applied to text-to-video diffusion models. To address these problem, we introduce an LLM-based prompt adaptation framework, termed as Prompt-A-Video, which excels in crafting Video-Centric, Labor-Free and Preference-Aligned prompts tailored to specific video diffusion model. Our approach involves a meticulously crafted two-stage optimization and alignment system. Initially, we conduct a reward-guided prompt evolution pipeline to automatically create optimal prompts pool and leverage them for supervised fine-tuning (SFT) of the LLM. Then multi-dimensional rewards are employed to generate pairwise data for the SFT model, followed by the direct preference optimization (DPO) algorithm to further facilitate preference alignment. Through extensive experimentation and comparative analyses, we validate the effectiveness of Prompt-A-Video across diverse generation models, highlighting its potential to push the boundaries of video generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee 等CVPR 2026 · 被引用 30 次
- FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video GenerationShilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge 等AAAI 2026 · 被引用 30 次
- A Systematic Survey of Automatic Prompt Optimization TechniquesKiran Ramnath, Kang Zhou, Sheng Guan, Soumya Smruti Mishra 等EMNLP 2025 · 被引用 5 次
- Diverse Video Generation with Determinantal Point Process-Guided Policy OptimizationTahira Kazimi, Connor Dunlop, Pinar YanardagCVPR 2026 · 被引用 4 次
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationJiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationBingjie Gao, Xinyu Gao, Xiaoxue Wu, Yujie Zhou 等CVPR 2025
- VideoDPO: Omni-Preference Alignment for Video Diffusion GenerationRuntao Liu, Haoyu Wu, Ziqiang Zheng, Chen Wei 等CVPR 2025
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 被引用 303 次
- InstructVideo: Instructing Video Diffusion Models with Human FeedbackHangjie Yuan, Shiwei Zhang, Xiang Wang, Yujie Wei 等CVPR 2024 · 被引用 13 次
- TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT AlignmentShicheng Li, Lei Li, Kun Ouyang, Shuhuai Ren 等AAAI 2026 · 被引用 2 次
