Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
Hongkang Li, Hancheng Min, Rene Vidal
摘要
Transformer-based diffusion models have demonstrated remarkable performance at generating high-quality samples. However, our theoretical understanding of the reasons for this success remains limited. For instance, existing models are typically trained by minimizing a denoising objective, which is equivalent to fitting the score function of the training data. However, we do not know why transformer-based models can match the score function for denoising, or why gradient-based methods converge to the optimal denoising model despite the non-convex loss landscape. To the best of our knowledge, this paper provides the first convergence analysis for training transformer-based diffusion models. More specifically, we consider the population Denoising Diffusion Probabilistic Model (DDPM) objective for denoising data that follow a multi-token Gaussian mixture distribution. We theoretically quantify the required number of tokens per data point and training iterations for the global convergence towards the Bayes optimal risk of the denoising objective, thereby achieving a desired score matching error. A deeper investigation reveals that the self-attention module of the trained transformer implements a mean denoising mechanism that enables the trained model to approximate the oracle Minimum Mean Squared Error (MMSE) estimator of the injected noise in the diffusion steps. Numerical experiments validate these findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- On the Generalization Properties of Diffusion ModelsPuheng Li, Zhong Li, Huishuai Zhang, Jiang BianNeurIPS 2023 · 被引用 86 次
相关 Paper
- O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal AssumptionsGen Li, Yuling YanICLR 2025 · 被引用 1 次
- Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion ModelsGen Li, Yuling YanNeurIPS 2024 · 被引用 66 次
- Towards Non-Asymptotic Convergence for Diffusion-Based Generative ModelsGen Li, Yuting Wei, Yuxin Chen, Yuejie ChiICLR 2024 · 被引用 39 次
- Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional SamplersYuchen Liang, Peizhong Ju, Yingbin Liang, Ness B. ShroffICLR 2025
- Assessing the quality of denoising diffusion models in Wasserstein distance: noisy score and optimal boundsVahan Arsenyan, Elen Vardanyan, Arnak S. DalalyanNeurIPS 2025 · 被引用 6 次
