Lune

EMNLP2024顶会

VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do

2024年份
20被引次数
58顶会引用

摘要

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metrics are able to provide reliable scores over generated videos. The main barrier is the lack of largescale human-annotated datasets. In this paper, we release VIDEOFEEDBACK, the first largescale dataset containing human-provided multiaspect score over 37.6K synthesized videos from 11 existing video generative models. We train VIDEOSCORE (initialized from Mantis) based on VIDEOFEEDBACK to enable automatic video quality assessment. Experiments show that the Spearman correlation between VIDEOSCORE and humans can reach 77.1 on VIDEOFEEDBACK-test, beating the prior best metrics by about 50 points. Further results on other held-out EvalCrafter, GenAI-Bench, and VBench show that VIDEOSCORE has consistently much higher correlation with human judges than other metrics. Due to these results, we believe VIDEOSCORE can serve as a great proxy for human raters to (1) rate different video models to track progress (2) simulate fine-grained human feedback in Reinforcement Learning with Human Feedback (RLHF) to improve current video generation models.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper58

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖