Ref4D-VideoBench: Four-Dimensional Reference-Based Evaluation of Text-to-Video Generative Models
Jiajia Wei, YuJia He, Yuhan Hou, Hang Qi, Sihua Wang, Jincheng Shi, Kwok Fung Li, Zibin Zheng, Weibin Wu
摘要
Most existing evaluations of generated videos adopt a no-reference paradigm. Although recent benchmarks cover multiple dimensions and show moderate correlation with human preferences, relying solely on textual prompts weakens real-world constraints and makes it difficult to produce accountable and interpretable judgments on instance-level issues such as target behavior deviation, temporal inconsistency, and commonsense violations. In scenarios with explicit expectations, such as controlled generation, reference videos naturally provide rich, unambiguous spatio-temporal evidence, enabling stricter and more trustworthy assessment. Motivated by this, we propose Ref4D, a reference-based, fine-grained, multi-dimensional benchmark for generated video evaluation. Ref4D contains 600 high-quality reference videos with tightly evidence-bounded prompts, and introduces a 12-metric structured evaluation suite along four key dimensions: basic semantic alignment, motion consistency, event temporal consistency, and world knowledge consistency. Experiments on eight text-to-video models show that Ref4D achieves stronger agreement with human judgments than representative no-reference frameworks, while precisely diagnosing the dimensions and causes of failure for each video. By integrating explicit reference evidence with multimodal reasoning, Ref4D provides a practical and human-aligned standard for generated video evaluation and a tool to guide the development of more reliable generative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen 等ICML 2024 · 被引用 499 次
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical PerspectivesHaoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen 等ICCV 2023 · 被引用 371 次
相关 Paper
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang 等CVPR 2024
- LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video GenerationXiangqing Zheng, CHENGYUE WU, Kehai Chen, Min zhangICML 2026 · 被引用 3 次
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsYiting Lu, Wei Luo, Peiyan Tu, Haoran Li 等CVPR 2026 · 被引用 10 次
- SVBench: Evaluation of Video Generation Models on Social ReasoningWenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li 等CVPR 2026 · 被引用 5 次
- GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step ReasoningZhun Mou, Bin Xia, Zhengchao Huang, Wenming Yang 等ICML 2025
