On the Content Bias in Fréchet Video Distance
Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu, Jia-Bin Huang
Abstract
a) Reference Videos (b) Medium Spatial & No Temporal Corruption (c) Small Spatial & Severe Temporal Corruption FVD=317.10 FVD=310.52 Figure 1. FVD is biased towards per-frame quality than temporal consistency. FVD [72], a commonly used video generation evaluation metric, should ideally capture both spatial and temporal aspects. However, our experiments reveal a strong bias toward individual frame quality. (b) First, we apply mild spatial distortions through local warping, which results in an FVD score of 317.10. (c) Next, we induce slightly less spatial corruptions but severe temporal inconsistencies by altering each frame differently. These changes create artifacts that are noticeable to humans and evident in the spatiotemporal x-t slice, as seen in the bottom row, but surprisingly lead to a lower FVD score of 310.52. This discrepancy highlights the metric's bias towards individual frame quality. We encourage readers to view the videos with Acrobat Reader or visit our website to observe the inconsistencies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- Vivid-ZOO: Multi-View Video Generation with Diffusion ModelBing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai et al.NeurIPS 2024 · 48 citations
- SF-V: Single Forward Video Generation ModelZhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu et al.NeurIPS 2024 · 43 citations
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion ModelsZiyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace et al.NeurIPS 2025 · 36 citations
- Fast and Memory-Efficient Video Diffusion Using Streamlined InferenceZheng Zhan, Yushu Wu, Yifan Gong, Zichong Meng et al.NeurIPS 2024 · 23 citations
- Self-Supervised Flow Matching for Scalable Multi-Modal SynthesisHila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell et al.ICML 2026 · 13 citations
Builds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative ModelsPum Jun Kim, Seojun Kim, Jaejun YooICLR 2024 · 11 citations
- Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution QualityGe Ya Luo, Gian Mario Favero, Zhi Hao Luo, Alexia Jolicoeur-Martineau et al.ICLR 2025
- Direct Motion Models for Assessing Generated VideosKelsey R. Allen, Carl Doersch, Guangyao Zhou, Mohammed Suhail et al.ICML 2025
- FovVideoVDP: a visible difference predictor for wide field-of-view videoRafal K. Mantiuk, Gyorgy Denes, Alexandre Chapiro, Anton Kaplanyan et al.SIGGRAPH 2021 · 158 citations
- EvalCrafter: Benchmarking and Evaluating Large Video Generation ModelsYaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang et al.CVPR 2024
