Lune

CVPR2024顶会

On the Content Bias in Fréchet Video Distance

Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu, Jia-Bin Huang

2024年份
34顶会引用

摘要

a) Reference Videos (b) Medium Spatial & No Temporal Corruption (c) Small Spatial & Severe Temporal Corruption FVD=317.10 FVD=310.52 Figure 1. FVD is biased towards per-frame quality than temporal consistency. FVD [72], a commonly used video generation evaluation metric, should ideally capture both spatial and temporal aspects. However, our experiments reveal a strong bias toward individual frame quality. (b) First, we apply mild spatial distortions through local warping, which results in an FVD score of 317.10. (c) Next, we induce slightly less spatial corruptions but severe temporal inconsistencies by altering each frame differently. These changes create artifacts that are noticeable to humans and evident in the spatiotemporal x-t slice, as seen in the bottom row, but surprisingly lead to a lower FVD score of 310.52. This discrepancy highlights the metric's bias towards individual frame quality. We encourage readers to view the videos with Acrobat Reader or visit our website to observe the inconsistencies.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper34

问问它们各自怎么用它

它引用的顶会 Paper37

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖