The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
Bingjie Gao, Xinyu Gao, Xiaoxue Wu, Yujie Zhou, Yu Qiao, Li Niu, Xinyuan Chen, Yaohui Wang
Abstract
The evolution of Text-to-video (T2V) generative models, trained on large-scale datasets, has been marked by significant progress. However, the sensitivity of T2V generative models to input prompts highlights the critical role of prompt design in influencing generative outcomes. Prior research has predominantly relied on Large Language Models (LLMs) to align user-provided prompts with the distribution of training prompts, albeit without tailored guidance encompassing prompt vocabulary and sentence structure nuances. To this end, we introduce RAPO, a novel Retrieval-Augmented Prompt Optimization framework. In order to address potential inaccuracies and ambiguous details generated by LLM-generated prompts. RAPO refines the naive prompts through dual optimization branches, selecting the superior prompt for T2V generation. The first branch augments user prompts with diverse modifiers extracted from a learned relational graph, refining them to align with the format of training prompts via a fine-tuned LLM. Conversely, the second branch rewrites the naive prompt using a pre-trained LLM following a well-defined instruction set. Extensive experiments demonstrate that RAPO can effectively enhance both the static and dynamic dimensions of generated videos, demonstrating the significance of prompt optimization for user-provided prompts. Project website: GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df754ddf-cf7c-49d1-9682-47e96d84b212Cited by top-tier papers10
- VISTA: A Test-Time Self-Improving Video Generation AgentDo Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee et al.CVPR 2026 · 30 citations
- LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text EncodersBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.NeurIPS 2025 · 17 citations
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMsYumin Choi, Dongki Kim, Jinheon Baek, Sung Ju HwangICLR 2026 · 4 citations
- Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual RepresentationBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.CVPR 2026 · 2 citations
- V.I.P.: Iterative Online Preference Distillation for Efficient Video Diffusion ModelsJisoo Kim, Wooseok Seo, Junwan Kim, Seungho Park et al.ICCV 2025 · 1 citation
Builds on33
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationJiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu et al.ICCV 2025 · 3 citations
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang et al.ICCV 2025 · 3 citations
- Multi-Modal Inductive Framework for Text-Video RetrievalQian Li, Yucheng Zhou, Cheng Ji, Feihong Lu et al.ACM MM 2024 · 8 citations
- VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal RetrievalSiteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang et al.CVPR 2023
- IPO: Interpretable Prompt Optimization for Vision-Language ModelsYingjun Du, Wenfang Sun, Cees SnoekNeurIPS 2024 · 15 citations
