Lune

CVPR2024Top-tier venue

On the Content Bias in Fréchet Video Distance

Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu, Jia-Bin Huang

2024Year
34Top-tier citations

Abstract

a) Reference Videos (b) Medium Spatial & No Temporal Corruption (c) Small Spatial & Severe Temporal Corruption FVD=317.10 FVD=310.52 Figure 1. FVD is biased towards per-frame quality than temporal consistency. FVD [72], a commonly used video generation evaluation metric, should ideally capture both spatial and temporal aspects. However, our experiments reveal a strong bias toward individual frame quality. (b) First, we apply mild spatial distortions through local warping, which results in an FVD score of 317.10. (c) Next, we induce slightly less spatial corruptions but severe temporal inconsistencies by altering each frame differently. These changes create artifacts that are noticeable to humans and evident in the spatiotemporal x-t slice, as seen in the bottom row, but surprisingly lead to a lower FVD score of 310.52. This discrepancy highlights the metric's bias towards individual frame quality. We encourage readers to view the videos with Acrobat Reader or visit our website to observe the inconsistencies.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers34

Ask how each one uses it

Builds on37

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines