Lune

NeurIPS2025Top-tier venue

An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models

Binxu Wang, Cengiz Pehlevan

2025Year
26Citations
11Top-tier citations

Abstract

We develop an analytical framework for understanding how the generated distribution evolves during diffusion model training. Leveraging a Gaussian-equivalence principle, we solve the full-batch gradient-flow dynamics of linear and convolutional denoisers and integrate the resulting probability-flow ODE, yielding analytic expressions for the generated distribution. The theory exposes a universal inverse-variance spectral law: the time for an eigen-or Fourier mode to match its target variance scales as τ ∝ λ -1 , so high-variance (coarse) structure is mastered orders of magnitude sooner than low-variance (fine) detail. Extending the analysis to deep linear networks and circulant full-width convolutions shows that weight sharing merely multiplies learning rates-accelerating but not eliminating the bias-whereas local convolution introduces a qualitatively different bias. Experiments on Gaussian and natural-image datasets confirm the spectral law persists in deep MLP-based UNet. Convolutional U-Nets, however, display rapid near-simultaneous emergence of many modes, implicating local convolution in reshaping learning dynamics. These results underscore how data covariance governs the order and speed with which diffusion models learn, and they call for deeper investigation of the unique inductive biases introduced by local convolution.

4 Learning in Diffusion Models with a Linear Denoiser Problem set-up. Throughout the paper, we assume the denoiser at each noise scale is linear (affine) and independent across scales:

Since the parameters W σ , b σ are decoupled across noise scales, each σ can be analysed independently. Through further parametrization, this umbrella form captures linear residual nets, deep linear nets, and linear convolutional nets (see Sec. 5).

We train on an arbitrary distribution p 0 with mean µ and covariance Σ by gradient flow on the full-batch DSM loss, i.e. the exact expectation over data and noise (2). (In practice, one cannot sample all z values, but the full-batch limit yields clean closed-form dynamics.)

This setting lets us dissect analytically the role of data spectrum, model architecture (W σ parametrisation), and loss variant in shaping diffusion learning.

Gaussian equivalence. For any joint distribution p(X, Y ) the quadratic loss

depends on p only through the first two moments of (X, Y ); see App. C.1.1 for proof. Hence a linear denoiser trained on arbitrary p 0 interacts with the data solely via its mean µ and covariance Σ.

Instance for diffusion. Under EDM loss (2), the noisy input-target pair is X = x 0 + σz, Y = x 0 , giving Σ XX = Σ + σ 2 I, Σ Y X = Σ.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4bb0bb3f-8ea0-41f2-a38e-c328c5c57923

Cited by top-tier papers11

Ask how each one uses it

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines