Analysis of Learning a Flow-based Generative Model from Limited Sample Complexity
Hugo Cui, Florent Krzakala, Eric Vanden-Eijnden, Lenka Zdeborová
Abstract
We study the problem of training a flow-based generative model, parametrized by a two-layer autoencoder, to sample from a high-dimensional Gaussian mixture. We provide a sharp end-to-end analysis of the problem. First, we provide a tight closed-form characterization of the learnt velocity field, when parametrized by a shallow denoising auto-encoder trained on a finite number n of samples from the target distribution. Building on this analysis, we provide a sharp description of the corresponding generative flow, which pushes the base Gaussian density forward to an approximation of the target density. In particular, we provide closed-form formulae for the distance between the means of the generated mixture and the mean of the target mixture, which we show decays as Θn( 1 /n). Finally, this rate is shown to be in fact Bayes-optimal. • We provide a sharp asymptotic closed-form characterization of the learnt velocity field, as a function of the target Gaussian mixture parameters, the stochastic interpolation schedule, and the number of training samples n. • We characterize the associated flow by providing a tight characterization of a small number of summary statistics, tracking the dynamics of a sample from the Gaussian base distribution as it is transported by the learnt velocity
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1006f1b-5a78-40c0-a50a-07a65773cf8bCited by top-tier papers16
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingTony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc MézardNeurIPS 2025 · 93 citations
- What does guidance do? A fine-grained analysis in a simple settingMuthu Chidambaram, Khashayar Gatmiry, Sitan Chen, Holden Lee et al.NeurIPS 2024 · 53 citations
- Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian StructureXiang Li, Yixiang Dai, Qing QuNeurIPS 2024 · 45 citations
- Critical windows: non-asymptotic theory for feature emergence in diffusion modelsMarvin Li, Sitan ChenICML 2024 · 34 citations
- Few-Shot Diffusion Models Escape the Curse of DimensionalityRuofeng Yang, Bo Jiang, Cheng Chen, Ruinan Jin et al.NeurIPS 2024 · 13 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- A solvable model of learning generative diffusion: theory and insightsHugo Cui, Cengiz Pehlevan, Yue M. LuNeurIPS 2025 · 11 citations
- High-dimensional Asymptotics of Denoising AutoencodersHugo Cui, Lenka ZdeborováNeurIPS 2023 · 26 citations
- Fundamental Limits of Two-layer Autoencoders, and Achieving Them with Gradient MethodsAleksandr Shevchenko, Kevin Kögler, Hamed Hassani, Marco MondelliICML 2023 · 3 citations
- Identifying through Flows for Recovering Latent RepresentationsShen Li, Bryan Hooi, Gim Hee LeeICLR 2020 · 15 citations
- Diffusion Normalizing FlowQinsheng Zhang, Yongxin ChenNeurIPS 2021 · 119 citations
