An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models
Binxu Wang, Cengiz Pehlevan
Abstract
We develop an analytical framework for understanding how the generated distribution evolves during diffusion model training. Leveraging a Gaussian-equivalence principle, we solve the full-batch gradient-flow dynamics of linear and convolutional denoisers and integrate the resulting probability-flow ODE, yielding analytic expressions for the generated distribution. The theory exposes a universal inverse-variance spectral law: the time for an eigen-or Fourier mode to match its target variance scales as τ ∝ λ -1 , so high-variance (coarse) structure is mastered orders of magnitude sooner than low-variance (fine) detail. Extending the analysis to deep linear networks and circulant full-width convolutions shows that weight sharing merely multiplies learning rates-accelerating but not eliminating the bias-whereas local convolution introduces a qualitatively different bias. Experiments on Gaussian and natural-image datasets confirm the spectral law persists in deep MLP-based UNet. Convolutional U-Nets, however, display rapid near-simultaneous emergence of many modes, implicating local convolution in reshaping learning dynamics. These results underscore how data covariance governs the order and speed with which diffusion models learn, and they call for deeper investigation of the unique inductive biases introduced by local convolution.
4 Learning in Diffusion Models with a Linear Denoiser Problem set-up. Throughout the paper, we assume the denoiser at each noise scale is linear (affine) and independent across scales:
Since the parameters W σ , b σ are decoupled across noise scales, each σ can be analysed independently. Through further parametrization, this umbrella form captures linear residual nets, deep linear nets, and linear convolutional nets (see Sec. 5).
We train on an arbitrary distribution p 0 with mean µ and covariance Σ by gradient flow on the full-batch DSM loss, i.e. the exact expectation over data and noise (2). (In practice, one cannot sample all z values, but the full-batch limit yields clean closed-form dynamics.)
This setting lets us dissect analytically the role of data spectrum, model architecture (W σ parametrisation), and loss variant in shaping diffusion learning.
Gaussian equivalence. For any joint distribution p(X, Y ) the quadratic loss
depends on p only through the first two moments of (X, Y ); see App. C.1.1 for proof. Hence a linear denoiser trained on arbitrary p 0 interacts with the data solely via its mean µ and covariance Σ.
Instance for diffusion. Under EDM loss (2), the noisy input-target pair is X = x 0 + σz, Y = x 0 , giving Σ XX = Σ + σ 2 I, Σ Y X = Σ.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bb0bb3f-8ea0-41f2-a38e-c328c5c57923Cited by top-tier papers11
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 15 citations
- A Random Matrix Perspective on the Consistency of Diffusion ModelsBinxu Wang, Jacob A Zavatone-Veth, Cengiz PehlevanICML 2026 · 4 citations
- Two Calm Ends and the Wild Middle: A Geometric Picture of Memorization in Diffusion ModelsNick Dodson, Xinyu Gao, Qingsong Wang, Yusu Wang et al.ICML 2026 · 4 citations
- Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMsHongkang Li, Hancheng Min, Rene VidalICML 2026 · 1 citation
- Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery LocalizationYiming Wang, Baiqi Wu, Qingming Li, Jiahao Chen et al.ICML 2026
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian StructureXiang Li, Yixiang Dai, Qing QuNeurIPS 2024 · 45 citations
- Locality in Image Diffusion Models Emerges from Data StatisticsArtem Lukoianov, Chenyang Yuan, Justin M. Solomon, Vincent SitzmannNeurIPS 2025 · 32 citations
- A theory of learning data statistics in diffusion models, from easy to hardLorenzo Bardone, Claudia Merger, Sebastian GoldtICML 2026
- Towards a Mechanistic Explanation of Diffusion Model GeneralizationMatthew Niedoba, Berend Zwartsenberg, Kevin Patrick Murphy, Frank WoodICML 2025
- On Inductive Biases That Enable Generalization in Diffusion TransformersJie An, De Wang, Pengsheng Guo, Jiebo Luo et al.NeurIPS 2025 · 1 citation
