VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao
Abstract
Multimodal large language models (MLLMs) have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation rely on MLLMs to evaluate the quality of generated videos, but the capabilities of MLLMs on interpreting AIGC videos remain largely underexplored. To address this, we propose a new benchmark, VF-EVAL, which introduces four tasks-coherence validation, error awareness, error type detection, and reasoning evaluation-to comprehensively evaluate the abilities of MLLMs on AIGC videos. We evaluate 13 frontier MLLMs on VF-EVAL and find that even the best-performing model, GPT-4.1, struggles to achieve consistently good performance across all tasks. This highlights the challenging nature of our benchmark. Additionally, to investigate the practical applications of VF-EVAL in improving video generation, we conduct an experiment, REPROMPT, demonstrating that aligning MLLMs more closely with human feedback can benefit video generation. Data songtingyu/VF-Eval Code SighingSnow/VF-Eval (a) Yes-Or-No (c) Open-Ended (b) Multichoice Q1: Is there moral issuse in this video, including human, meaningless text, violence ? Q2: Is there distortion issue within the pink pig toy? A. Yes B. No A. Yes B. No Q: What is unusual about the straw's appearance? A. The straw is missing its top part. B. The colors of the top and bottom parts are different. C. The straw is shorter than a regular one. D. The straw is bent at an unusual angle. Q1: Identify any discrepancies between the video content and "A soccer player kicks a ball harder, making it travel farther than a light tap. " Afterward, suggest a better prompt based on the text to help regenerate the video. A1: (1) Mis-alignment: There are two soccer balls in the video, and the soccer player does not kick the ball out. (2) Better prompt: A soccer player kicks a soccer ball hard. Q2: How many soccer balls does the man kick? A2: The man in the video actually kicks one ball, but the trajectory of the ball he kicked does not match his action, while the trajectory of the other soccer ball does. And a kick on one ball can't make two balls move.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ef78d5c-34ad-4e10-80a2-16bca224aa0aCited by top-tier papers2
- Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative ModelsJiajia Wei, YuJia He, Yuhan Hou, Hang Qi et al.CVPR 2026
- Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content EvaluationTianjiao Li, Kai Zhao, Xiang Li, Yang Liu et al.ACL 2026
Builds on13
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchySimon Ging, María Alejandra Bravo, Thomas BroxICLR 2024 · 24 citations
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou et al.EMNLP 2023 · 3 citations
Related papers
- Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsZicheng Zhang, Ziheng Jia, Haoning Wu, Chunyi Li et al.CVPR 2025
- VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement LearningXuanyu Zhang, Weiqi Li, Shijie Zhao, Junlin Li et al.AAAI 2026 · 20 citations
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang et al.ICML 2026 · 14 citations
- Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMsZijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du et al.ICLR 2025
- Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQAYue Fan, Jing Gu, Kaiwen Zhou, Qianqi Yan et al.ACL 2024 · 1 citation
