VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao
摘要
Multimodal large language models (MLLMs) have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation rely on MLLMs to evaluate the quality of generated videos, but the capabilities of MLLMs on interpreting AIGC videos remain largely underexplored. To address this, we propose a new benchmark, VF-EVAL, which introduces four tasks-coherence validation, error awareness, error type detection, and reasoning evaluation-to comprehensively evaluate the abilities of MLLMs on AIGC videos. We evaluate 13 frontier MLLMs on VF-EVAL and find that even the best-performing model, GPT-4.1, struggles to achieve consistently good performance across all tasks. This highlights the challenging nature of our benchmark. Additionally, to investigate the practical applications of VF-EVAL in improving video generation, we conduct an experiment, REPROMPT, demonstrating that aligning MLLMs more closely with human feedback can benefit video generation. Data songtingyu/VF-Eval Code SighingSnow/VF-Eval (a) Yes-Or-No (c) Open-Ended (b) Multichoice Q1: Is there moral issuse in this video, including human, meaningless text, violence ? Q2: Is there distortion issue within the pink pig toy? A. Yes B. No A. Yes B. No Q: What is unusual about the straw's appearance? A. The straw is missing its top part. B. The colors of the top and bottom parts are different. C. The straw is shorter than a regular one. D. The straw is bent at an unusual angle. Q1: Identify any discrepancies between the video content and "A soccer player kicks a ball harder, making it travel farther than a light tap. " Afterward, suggest a better prompt based on the text to help regenerate the video. A1: (1) Mis-alignment: There are two soccer balls in the video, and the soccer player does not kick the ball out. (2) Better prompt: A soccer player kicks a soccer ball hard. Q2: How many soccer balls does the man kick? A2: The man in the video actually kicks one ball, but the trajectory of the ball he kicked does not match his action, while the trajectory of the other soccer ball does. And a kick on one ball can't make two balls move.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative ModelsJiajia Wei, YuJia He, Yuhan Hou, Hang Qi 等CVPR 2026
- Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content EvaluationTianjiao Li, Kai Zhao, Xiang Li, Yang Liu 等ACL 2026
它引用的顶会 Paper13
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 等ICML 2024 · 被引用 464 次
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun 等CVPR 2024 · 被引用 83 次
- Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchySimon Ging, María Alejandra Bravo, Thomas BroxICLR 2024 · 被引用 24 次
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou 等EMNLP 2023 · 被引用 3 次
相关 Paper
- Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsZicheng Zhang, Ziheng Jia, Haoning Wu, Chunyi Li 等CVPR 2025
- VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement LearningXuanyu Zhang, Weiqi Li, Shijie Zhao, Junlin Li 等AAAI 2026 · 被引用 20 次
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang 等ICML 2026 · 被引用 14 次
- Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMsZijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 等ICLR 2025
- Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQAYue Fan, Jing Gu, Kaiwen Zhou, Qianqi Yan 等ACL 2024 · 被引用 1 次
