Divergence Frontiers for Generative Models: Sample Complexity, Quantization Effects, and Frontier Integrals
Lang Liu, Krishna Pillutla, Sean Welleck, Sewoong Oh, Yejin Choi, Zaïd Harchaoui
Abstract
The spectacular success of deep generative models calls for quantitative tools to measure their statistical performance. Divergence frontiers have recently been proposed as an evaluation framework for generative models, due to their ability to measure the quality-diversity trade-off inherent to deep generative modeling. We establish non-asymptotic bounds on the sample complexity of divergence frontiers. We also introduce frontier integrals which provide summary statistics of divergence frontiers. We show how smoothed estimators such as Good-Turing or Krichevsky-Trofimov can overcome the missing mass problem and lead to faster rates of convergence. We illustrate the theoretical results with numerical examples from natural language processing and computer vision. While this framework is mathematically elegant and empirically successful [37, 49] , the statistical properties of divergence frontiers are not well understood. Estimating divergence frontiers from data for large generative models involves two approximations: (a) joint quantization of the model distribution and the target distribution into discrete distributions with quantization level k, and (b) statistical estimation of the divergence frontiers based on the quantized distributions. Djolonga et al. [18] argue that the quantization often introduces a positive bias, making the distributions appear closer than they really are; while a small sample size can result in a pessimistic estimate of the divergence frontier. The latter effect is due to the missing mass of the samples, causing the two distributions to appear farther than they really are because the samples do not cover some parts of the distributions. The first consideration favors a large k, while the second favors a small k. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db53a6ae-a87f-4075-b52f-2b1b85da80eaCited by top-tier papers14
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Remasking Discrete Diffusion Models with Inference-Time ScalingGuanghan Wang, Yair Schiff, Subham S. Sahoo, Volodymyr KuleshovNeurIPS 2025 · 199 citations
- Does Writing with Language Models Reduce Content Diversity?Vishakh Padmakumar, He HeICLR 2024 · 173 citations
- DistiLLM: Towards Streamlined Distillation for Large Language ModelsJongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young YunICML 2024 · 86 citations
- Continuously Augmented Discrete Diffusion model for Categorical Generative ModelingHuangjie Zheng, Shansan Gong, Ruixiang Zhang, Tianrong Chen et al.ICLR 2026 · 30 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi et al.ICML 2020 · 553 citations
Related papers
- A Theoretical Framework for Statistical Evaluability of Generative ModelsShashaank Aiyer, Yishay Mansour, Shay Moran, Han ShaoICML 2026 · 1 citation
- Feature Likelihood Score: Evaluating the Generalization of Generative Models Using SamplesMarco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin et al.NeurIPS 2023 · 39 citations
- Evaluating Lossy Compression Rates of Deep Generative ModelsSicong Huang, Alireza Makhzani, Yanshuai Cao, Roger B. GrosseICML 2020 · 30 citations
- Understanding Deep Generative Models with Generalized Empirical LikelihoodsSuman V. Ravuri, Mélanie Rey, Shakir Mohamed, Marc Peter DeisenrothCVPR 2023
- Theoretical guarantees on the best-of-n alignment policyAhmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander Nicholas D'Amour et al.ICML 2025
