Feature Likelihood Score: Evaluating the Generalization of Generative Models Using Samples
Marco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin, Yoram Bachrach, Gauthier Gidel
摘要
The past few years have seen impressive progress in the development of deep generative models capable of producing high-dimensional, complex, and photo-realistic data. However, current methods for evaluating such models remain incomplete: standard likelihood-based metrics do not always apply and rarely correlate with perceptual fidelity, while sample-based metrics, such as FID, are insensitive to overfitting, i.e., inability to generalize beyond the training set. To address these limitations, we propose a new metric called the Feature Likelihood Divergence (FLD), a parametric sample-based metric that uses density estimation to provide a comprehensive trichotomic evaluation accounting for novelty (i.e., different from the training samples), fidelity, and diversity of generated samples. We empirically demonstrate the ability of FLD to identify overfitting problem cases, even when previously proposed metrics fail. We also extensively evaluate FLD on various image datasets and model classes, demonstrating its ability to match intuitions of previous metrics like FID while offering a more comprehensive evaluation of generative models. Code is available at https://github.com/marcojira/fld .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Towards a Scalable Reference-Free Evaluation of Generative ModelsAzim Ospanov, Jingwei Zhang, Mohammad Jalali, Xuenan Cao 等NeurIPS 2024 · 被引用 32 次
- An Interpretable Evaluation of Entropy-based Novelty of Generative ModelsJingwei Zhang, Cheuk Ting Li, Farzan FarniaICML 2024 · 被引用 18 次
- SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE ScoreMohammad Jalali, Haoyu Lei, Amin Gohari, Farzan FarniaNeurIPS 2025 · 被引用 15 次
- Glauber Generative Model: Discrete Diffusion Models via Binary ClassificationHarshit Varma, Dheeraj Mysore Nagaraj, Karthikeyan ShanmugamICLR 2025
- Unveiling Differences in Generative Models: A Scalable Differential Clustering ApproachJingwei Zhang, Mohammad Jalali, Cheuk Ting Li, Farzan FarniaCVPR 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Deep Models Under the GAN: Information Leakage from Collaborative Deep LearningBriland Hitaj, Giuseppe Ateniese, Fernando Pérez-CruzCCS 2017 · 被引用 1,581 次
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 被引用 958 次
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 被引用 287 次
相关 Paper
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi 等ICML 2020 · 被引用 553 次
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner 等CVPR 2024
- Understanding Deep Generative Models with Generalized Empirical LikelihoodsSuman V. Ravuri, Mélanie Rey, Shakir Mohamed, Marc Peter DeisenrothCVPR 2023
- A Unifying Information-theoretic Perspective on Evaluating Generative ModelsAlexis Fox, Samarth Swarup, Abhijin AdigaAAAI 2025 · 被引用 1 次
- Rarity Score : A New Metric to Evaluate the Uncommonness of Synthesized ImagesJiyeon Han, Hwanil Choi, Yunjey Choi, Junho Kim 等ICLR 2023 · 被引用 7 次
