The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
Bingjie Gao, Xinyu Gao, Xiaoxue Wu, Yujie Zhou, Yu Qiao, Li Niu, Xinyuan Chen, Yaohui Wang
摘要
The evolution of Text-to-video (T2V) generative models, trained on large-scale datasets, has been marked by significant progress. However, the sensitivity of T2V generative models to input prompts highlights the critical role of prompt design in influencing generative outcomes. Prior research has predominantly relied on Large Language Models (LLMs) to align user-provided prompts with the distribution of training prompts, albeit without tailored guidance encompassing prompt vocabulary and sentence structure nuances. To this end, we introduce RAPO, a novel Retrieval-Augmented Prompt Optimization framework. In order to address potential inaccuracies and ambiguous details generated by LLM-generated prompts. RAPO refines the naive prompts through dual optimization branches, selecting the superior prompt for T2V generation. The first branch augments user prompts with diverse modifiers extracted from a learned relational graph, refining them to align with the format of training prompts via a fine-tuned LLM. Conversely, the second branch rewrites the naive prompt using a pre-trained LLM following a well-defined instruction set. Extensive experiments demonstrate that RAPO can effectively enhance both the static and dynamic dimensions of generated videos, demonstrating the significance of prompt optimization for user-provided prompts. Project website: GitHub.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee 等CVPR 2026 · 被引用 30 次
- LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text EncodersBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang 等NeurIPS 2025 · 被引用 17 次
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMsYumin Choi, Dongki Kim, Jinheon Baek, Sung Ju HwangICLR 2026 · 被引用 4 次
- Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual RepresentationBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang 等CVPR 2026 · 被引用 2 次
- V.I.P.: Iterative Online Preference Distillation for Efficient Video Diffusion ModelsJisoo Kim, Wooseok Seo, Junwan Kim, Seungho Park 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper33
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationJiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu 等ICCV 2025 · 被引用 3 次
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang 等ICCV 2025 · 被引用 3 次
- Multi-Modal Inductive Framework for Text-Video RetrievalQian Li, Yucheng Zhou, Cheng Ji, Feihong Lu 等ACM MM 2024 · 被引用 8 次
- VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal RetrievalSiteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang 等CVPR 2023
- IPO: Interpretable Prompt Optimization for Vision-Language ModelsYingjun Du, Wenfang Sun, Cees SnoekNeurIPS 2024 · 被引用 15 次
