VISTA: A Test-Time Self-Improving Video Generation Agent
Do Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee, Tomas Pfister, Sercan O Arik
Abstract
Despite rapid advances in text-to-video synthesis, generated video quality remains critically dependent on precise user prompts. Existing test-time optimization methods, successful in other domains, struggle with the multi-faceted nature of video. In this work, we introduce VISTA, a novel multi-agent system that autonomously improves video generation through refining prompts in an iterative loop. VISTA first decomposes a user's idea into a structured temporal plan. After generation, the best video is identified through a robust pairwise tournament. This winning video is then critiqued by a trio of specialized agents focusing on visual, audio, and contextual fidelity. Finally, a reasoning agent synthesizes this feedback to introspectively rewrite and enhance the prompt for the next generation cycle. Experiments on single-and multi-scene video generation scenarios show that while prior methods yield inconsistent gains, VISTA consistently improves video quality and alignment with user intent, achieving up to 60% pairwise win rate against state-of-the-art baselines. Human evaluators concur, preferring VISTA's outputs in 66.4% of comparisons. Direct Prompting (DP) VISTA (Ours) Single-scene (Polyak et al., 2025): The person's forehead creased with worry as he listened to bad news. Single-scene (Polyak et al., 2025): A spaceship entering hyperdrive, stars streaking past as it accelerates. Multi-scene-Interviews: The video features a man outdoors, asking a trivia question about a comedian known. . . Multi-scene-Animation: The video is an educational animation designed for young children. . .
Table 1 | Example videos generated by VISTA, showing improvements in visual fidelity, camera focus, smooth transitions, compelling storylines, and sounds. Optimized prompts and more examples are in our project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 303 citations
Related papers
- V-Stylist: Video Stylization via Collaboration and Reflection of MLLM AgentsZhengrong Yue, Shaobin Zhuang, Kunchang Li, Yanbo Ding et al.CVPR 2025
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationJiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu et al.ICCV 2025 · 3 citations
- The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationBingjie Gao, Xinyu Gao, Xiaoxue Wu, Yujie Zhou et al.CVPR 2025
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang et al.ICCV 2025 · 3 citations
- What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific PresentationsDongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon et al.ACL 2025
