Effectively Unbiased FID and Inception Score and Where to Find Them
Min Jin Chong, David A. Forsyth
Abstract
This paper shows that two commonly used evaluation metrics for generative models, the Fréchet Inception Distance (FID) and the Inception Score (IS), are biased -the expected value of the score computed for a finite sample set is not the true value of the score. Worse, the paper shows that the bias term depends on the particular model being evaluated, so model A may get a better score than model B simply because model A's bias term is smaller. This effect cannot be fixed by evaluating at a fixed number of samples. This means all comparisons using FID or IS as currently computed are unreliable.
We then show how to extrapolate the score to obtain an effectively bias-free estimate of scores computed with an infinite number of samples, which we term FID ∞ and IS ∞ . In turn, this effectively bias-free estimate requires good estimates of scores with a finite number of samples. We show that using Quasi-Monte Carlo integration notably improves estimates of FID and IS for finite sample sets. Our extrapolated scores are simple, drop-in replacements for the finite sample scores. Additionally, we show that using low discrepancy sequence in GAN training offers small improvements in the resulting generator. The code for calculating FID ∞ and IS ∞ is at https://github.com/ mchong6/FID_IS_infinity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83f4e74c-195c-44c0-9fe6-7a58cc664c68Cited by top-tier papers40
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth et al.NeurIPS 2022 · 496 citations
- Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion modelsGeorge Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui et al.NeurIPS 2023 · 260 citations
- On Aliased Resizing and Surprising Subtleties in GAN EvaluationGaurav Parmar, Richard Zhang, Jun-Yan ZhuCVPR 2022 · 250 citations
- Visual Fourier Prompt TuningRunjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu et al.NeurIPS 2024 · 58 citations
- ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12Liuqing Chen, Shuhong Xiao, Yunnong Chen, Yaxuan Song et al.CHI 2024 · 56 citations
Builds on1
Related papers
- Rethinking FID: Towards a Better Evaluation Metric for Image GenerationSadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner et al.CVPR 2024
- The Role of ImageNet Classes in Fréchet Inception DistanceTuomas Kynkäänniemi, Tero Karras, Miika Aittala, Timo Aila et al.ICLR 2023 · 44 citations
- TopP&R: Robust Support Estimation Approach for Evaluating Fidelity and Diversity in Generative ModelsPum Jun Kim, Yoojin Jang, Jisu Kim, Jaejun YooNeurIPS 2023 · 16 citations
- PSA-GAN: Progressive Self Attention GANs for Synthetic Time SeriesPaul Jeha, Michael Bohlke-Schneider, Pedro Mercado, Shubham Kapoor et al.ICLR 2022 · 92 citations
- Influence Estimation for Generative Adversarial NetworksNaoyuki Terashita, Hiroki Ohashi, Yuichi Nonaka, Takashi KanemaruICLR 2021 · 12 citations
