Lune

NeurIPS2025顶会

An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models

Binxu Wang, Cengiz Pehlevan

2025年份
26被引次数
11顶会引用

摘要

We develop an analytical framework for understanding how the generated distribution evolves during diffusion model training. Leveraging a Gaussian-equivalence principle, we solve the full-batch gradient-flow dynamics of linear and convolutional denoisers and integrate the resulting probability-flow ODE, yielding analytic expressions for the generated distribution. The theory exposes a universal inverse-variance spectral law: the time for an eigen-or Fourier mode to match its target variance scales as τ ∝ λ -1 , so high-variance (coarse) structure is mastered orders of magnitude sooner than low-variance (fine) detail. Extending the analysis to deep linear networks and circulant full-width convolutions shows that weight sharing merely multiplies learning rates-accelerating but not eliminating the bias-whereas local convolution introduces a qualitatively different bias. Experiments on Gaussian and natural-image datasets confirm the spectral law persists in deep MLP-based UNet. Convolutional U-Nets, however, display rapid near-simultaneous emergence of many modes, implicating local convolution in reshaping learning dynamics. These results underscore how data covariance governs the order and speed with which diffusion models learn, and they call for deeper investigation of the unique inductive biases introduced by local convolution.

4 Learning in Diffusion Models with a Linear Denoiser Problem set-up. Throughout the paper, we assume the denoiser at each noise scale is linear (affine) and independent across scales:

Since the parameters W σ , b σ are decoupled across noise scales, each σ can be analysed independently. Through further parametrization, this umbrella form captures linear residual nets, deep linear nets, and linear convolutional nets (see Sec. 5).

We train on an arbitrary distribution p 0 with mean µ and covariance Σ by gradient flow on the full-batch DSM loss, i.e. the exact expectation over data and noise (2). (In practice, one cannot sample all z values, but the full-batch limit yields clean closed-form dynamics.)

This setting lets us dissect analytically the role of data spectrum, model architecture (W σ parametrisation), and loss variant in shaping diffusion learning.

Gaussian equivalence. For any joint distribution p(X, Y ) the quadratic loss

depends on p only through the first two moments of (X, Y ); see App. C.1.1 for proof. Hence a linear denoiser trained on arbitrary p 0 interacts with the data solely via its mean µ and covariance Σ.

Instance for diffusion. Under EDM loss (2), the noisy input-target pair is X = x 0 + σz, Y = x 0 , giving Σ XX = Σ + σ 2 I, Σ Y X = Σ.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper11

问问它们各自怎么用它

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖