Statistically Optimal Generative Modeling with Maximum Deviation from the Empirical Distribution
Elen Vardanyan, Sona Hunanyan, Tigran Galstyan, Arshak Minasyan, Arnak S. Dalalyan
Abstract
This paper explores the problem of generative modeling, aiming to simulate diverse examples from an unknown distribution based on observed examples. While recent studies have focused on quantifying the statistical precision of popular algorithms, there is a lack of mathematical evaluation regarding the non-replication of observed examples and the creativity of the generative model. We present theoretical insights into this aspect, demonstrating that the Wasserstein GAN, constrained to left-invertible push-forward maps, generates distributions that not only avoid replication but also significantly deviate from the empirical distribution. Importantly, we show that leftinvertibility achieves this without compromising the statistical optimality of the resulting generator. Our most important contribution provides a finitesample lower bound on the Wasserstein-1 distance between the generative distribution and the empirical one. We also establish a finite-sample upper bound on the distance between the generative distribution and the true data-generating one. Both bounds are explicit and show the impact of key parameters such as sample size, dimensions of the ambient and latent spaces, noise level, and smoothness measured by the Lipschitz constant.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d17fde54-0e4c-46de-8ee0-9b1eb6c21015Builds on13
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Diffusion Schrödinger Bridge with Applications to Score-Based Generative ModelingValentin De Bortoli, James Thornton, Jeremy Heng, Arnaud DoucetNeurIPS 2021 · 811 citations
- Riemannian Score-Based Generative ModellingValentin De Bortoli, Emile Mathieu, Michael J. Hutchinson, James Thornton et al.NeurIPS 2022 · 306 citations
- Understanding and Mitigating Copying in Diffusion ModelsGowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping et al.NeurIPS 2023 · 265 citations
- Exactly Computing the Local Lipschitz Constant of ReLU NetworksMatt Jordan, Alexandros G. DimakisNeurIPS 2020 · 156 citations
Related papers
- Minimax Optimality (Probably) Doesn't Imply Distribution Learning for GANsSitan Chen, Jerry Li, Yuanzhi Li, Raghu MekaICLR 2022 · 6 citations
- Can Push-forward Generative Models Fit Multimodal Distributions?Antoine Salmona, Valentin De Bortoli, Julie Delon, Agnès DesolneuxNeurIPS 2022 · 53 citations
- Adversarial Lipschitz RegularizationDávid TerjékICLR 2020 · 55 citations
- Non-asymptotic Error Bounds for Bidirectional GANsShiao Liu, Yunfei Yang, Jian Huang, Yuling Jiao et al.NeurIPS 2021 · 8 citations
- Towards Generalized Implementation of Wasserstein Distance in GANsMinkai XuAAAI 2021 · 15 citations
