Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution Quality
Ge Ya Luo, Gian Mario Favero, Zhi Hao Luo, Alexia Jolicoeur-Martineau, Christopher Pal
摘要
The Fréchet Video Distance (FVD) is a widely adopted metric for evaluating video generation distribution quality. However, its effectiveness relies on critical assumptions. Our analysis reveals three significant limitations: (1) the non-Gaussianity of the Inflated 3D ConvNet (I3D) feature space; (2) the insensitivity of I3D features to temporal distortions; (3) the impractical sample sizes required for reliable estimation. These findings undermine FVD's reliability and show that FVD falls short as a standalone metric for video generation evaluation. After extensive analysis of a wide range of metrics and backbone architectures, we propose JEDi, the JEPA Embedding Distance, based on features derived from a Joint Embedding Predictive Architecture, measured using Maximum Mean Discrepancy with polynomial kernel. Our experiments on multiple open-source datasets show clear evidence that it is a superior alternative to the widely used FVD metric, requiring only 16% of the samples to reach its steady value, while increasing alignment with human evaluation by 34%, on average. Project page: https://oooolga.github.io/JEDi.github.io/ . Published as a conference paper at ICLR 2025 2. JEDi significantly reduces the number of samples needed to make an accurate estimate by using an MMD metric in a V-JEPA feature space, enabling reliable use in smaller datasets that do not meet the requirement when using FVD. 3. JEDi leverages the robust representations of a V-JEPA model, which are found to be more aligned with human evaluations compared to FVD. BACKGROUND AND NOTATIONS 2.1 VIDEO FEATURE REPRESENTATION Inflated 3D ConvNet: The Inflated 3D ConvNet (I3D) (Carreira & Zisserman, 2018) is a convolutional neural network model based on the pre-trained Inception-v1. It extends the 2D convolutional filters to 3D by replicating them along the temporal dimension. I3D, pre-trained on Kinetics, has demonstrated excellent classification performance on UCF-101 (Soomro et al., 2012), HMDB-51 (Kuehne et al., 2011), and Kinetics datasets (Kay et al., 2017), proving to be a valuable network for video recognition tasks. The original FVD work by Unterthiner et al. (2019) explores the use of I3D features trained on the Kinetics datasets. They analyze the features from the logits layer, as well as the features from the last pooling layer trained on the Kinetics-400 and Kinetics-600 datasets. Their experiments suggest that the features from the logits layer trained on the Kinetics-400 dataset are the most suitable for the FVD metric.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Inference-time Physics Alignment of Video Generative Models with Latent World ModelsJianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez 等CVPR 2026 · 被引用 32 次
- VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied AgentsGeorge Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen 等CVPR 2026
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei 等ICCV 2023 · 被引用 1,113 次
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 被引用 434 次
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin 等ICLR 2023 · 被引用 313 次
相关 Paper
- STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative ModelsPum Jun Kim, Seojun Kim, Jaejun YooICLR 2024 · 被引用 11 次
- Direct Motion Models for Assessing Generated VideosKelsey R. Allen, Carl Doersch, Guangyao Zhou, Mohammed Suhail 等ICML 2025
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner 等CVPR 2024
- On the Content Bias in Fréchet Video DistanceSongwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu 等CVPR 2024
- Fréchet Wavelet Distance: A Domain-Agnostic Metric for Image GenerationLokesh Veeramacheneni, Moritz Wolter, Hilde Kuehne, Juergen GallICLR 2025
