Analysis of Learning a Flow-based Generative Model from Limited Sample Complexity
Hugo Cui, Florent Krzakala, Eric Vanden-Eijnden, Lenka Zdeborová
摘要
We study the problem of training a flow-based generative model, parametrized by a two-layer autoencoder, to sample from a high-dimensional Gaussian mixture. We provide a sharp end-to-end analysis of the problem. First, we provide a tight closed-form characterization of the learnt velocity field, when parametrized by a shallow denoising auto-encoder trained on a finite number n of samples from the target distribution. Building on this analysis, we provide a sharp description of the corresponding generative flow, which pushes the base Gaussian density forward to an approximation of the target density. In particular, we provide closed-form formulae for the distance between the means of the generated mixture and the mean of the target mixture, which we show decays as Θn( 1 /n). Finally, this rate is shown to be in fact Bayes-optimal. • We provide a sharp asymptotic closed-form characterization of the learnt velocity field, as a function of the target Gaussian mixture parameters, the stochastic interpolation schedule, and the number of training samples n. • We characterize the associated flow by providing a tight characterization of a small number of summary statistics, tracking the dynamics of a sample from the Gaussian base distribution as it is transported by the learnt velocity
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingTony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc MézardNeurIPS 2025 · 被引用 93 次
- What does guidance do? A fine-grained analysis in a simple settingMuthu Chidambaram, Khashayar Gatmiry, Sitan Chen, Holden Lee 等NeurIPS 2024 · 被引用 53 次
- Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian StructureXiang Li, Yixiang Dai, Qing QuNeurIPS 2024 · 被引用 45 次
- Critical windows: non-asymptotic theory for feature emergence in diffusion modelsMarvin Li, Sitan ChenICML 2024 · 被引用 34 次
- Few-Shot Diffusion Models Escape the Curse of DimensionalityRuofeng Yang, Bo Jiang, Cheng Chen, Ruinan Jin 等NeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
相关 Paper
- A solvable model of learning generative diffusion: theory and insightsHugo Cui, Cengiz Pehlevan, Yue M. LuNeurIPS 2025 · 被引用 11 次
- High-dimensional Asymptotics of Denoising AutoencodersHugo Cui, Lenka ZdeborováNeurIPS 2023 · 被引用 26 次
- Fundamental Limits of Two-layer Autoencoders, and Achieving Them with Gradient MethodsAleksandr Shevchenko, Kevin Kögler, Hamed Hassani, Marco MondelliICML 2023 · 被引用 3 次
- Identifying through Flows for Recovering Latent RepresentationsShen Li, Bryan Hooi, Gim Hee LeeICLR 2020 · 被引用 15 次
- Diffusion Normalizing FlowQinsheng Zhang, Yongxin ChenNeurIPS 2021 · 被引用 119 次
