Reliable Fidelity and Diversity Metrics for Generative Models
Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, Jaejun Yoo
摘要
Devising indicative evaluation metrics for the image generation task remains an open problem. The most widely used metric for measuring the similarity between real and generated images has been the Fréchet Inception Distance (FID) score. Because it does not differentiate the fidelity and diversity aspects of the generated images, recent papers have introduced variants of precision and recall metrics to diagnose those properties separately. In this paper, we show that even the latest version of the precision and recall metrics are not reliable yet. For example, they fail to detect the match between two identical distributions, they are not robust against outliers, and the evaluation hyperparameters are selected arbitrarily. We propose density and coverage metrics that solve the above issues. We analytically and experimentally show that density and coverage provide more interpretable and reliable signals for practitioners than the existing metrics. Code: github.com/clovaai/generative-evaluation-prdc .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper164
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- StyleGAN-XL: Scaling StyleGAN to Large Diverse DatasetsAxel Sauer, Katja Schwarz, Andreas GeigerSIGGRAPH 2022 · 被引用 326 次
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 被引用 287 次
- Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion modelsGeorge Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui 等NeurIPS 2023 · 被引用 260 次
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold 等ICLR 2024 · 被引用 173 次
相关 Paper
- Probabilistic Precision and Recall Towards Reliable Evaluation of Generative ModelsDogyun Park, Suhyun KimICCV 2023 · 被引用 12 次
- Enhanced Generative Model Evaluation with Clipped Density and CoverageNicolas Salvy, Hugues Talbot, Bertrand ThirionICLR 2026 · 被引用 3 次
- Emergent Asymmetry of Precision and Recall for Measuring Fidelity and Diversity of Generative Models in High DimensionsMahyar Khayatkhoei, Wael Abd-AlmageedICML 2023 · 被引用 11 次
- Feature Likelihood Score: Evaluating the Generalization of Generative Models Using SamplesMarco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin 等NeurIPS 2023 · 被引用 39 次
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner 等CVPR 2024
