Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
Hongkang Li, Hancheng Min, Rene Vidal
Abstract
Transformer-based diffusion models have demonstrated remarkable performance at generating high-quality samples. However, our theoretical understanding of the reasons for this success remains limited. For instance, existing models are typically trained by minimizing a denoising objective, which is equivalent to fitting the score function of the training data. However, we do not know why transformer-based models can match the score function for denoising, or why gradient-based methods converge to the optimal denoising model despite the non-convex loss landscape. To the best of our knowledge, this paper provides the first convergence analysis for training transformer-based diffusion models. More specifically, we consider the population Denoising Diffusion Probabilistic Model (DDPM) objective for denoising data that follow a multi-token Gaussian mixture distribution. We theoretically quantify the required number of tokens per data point and training iterations for the global convergence towards the Bayes optimal risk of the denoising objective, thereby achieving a desired score matching error. A deeper investigation reveals that the self-attention module of the trained transformer implements a mean denoising mechanism that enables the trained model to approximate the oracle Minimum Mean Squared Error (MMSE) estimator of the injected noise in the diffusion steps. Numerical experiments validate these findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca69e5be-f160-414d-b18b-55849d22d8b0Builds on22
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- On the Generalization Properties of Diffusion ModelsPuheng Li, Zhong Li, Huishuai Zhang, Jiang BianNeurIPS 2023 · 86 citations
Related papers
- O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal AssumptionsGen Li, Yuling YanICLR 2025 · 1 citation
- Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion ModelsGen Li, Yuling YanNeurIPS 2024 · 66 citations
- Towards Non-Asymptotic Convergence for Diffusion-Based Generative ModelsGen Li, Yuting Wei, Yuxin Chen, Yuejie ChiICLR 2024 · 39 citations
- Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional SamplersYuchen Liang, Peizhong Ju, Yingbin Liang, Ness B. ShroffICLR 2025
- Assessing the quality of denoising diffusion models in Wasserstein distance: noisy score and optimal boundsVahan Arsenyan, Elen Vardanyan, Arnak S. DalalyanNeurIPS 2025 · 6 citations
