Lune

NeurIPS2025顶会

Locality in Image Diffusion Models Emerges from Data Statistics

Artem Lukoianov, Chenyang Yuan, Justin M. Solomon, Vincent Sitzmann

2025年份
32被引次数
11顶会引用

摘要

Recent work has shown that the generalization ability of image diffusion models arises from the locality properties of the trained neural network. In particular, when denoising a particular pixel, the model relies on a limited neighborhood of the input image around that pixel, which, according to the previous work, is tightly related to the ability of these models to produce novel images. Since locality is central to generalization, it is crucial to understand why diffusion models learn local behavior in the first place, as well as the factors that govern the properties of locality patterns. In this work, we present evidence that the locality in deep diffusion models emerges as a statistical property of the image dataset and is not due to the inductive bias of convolutional neural networks, as suggested in previous work. Specifically, we demonstrate that an optimal parametric linear denoiser exhibits similar locality properties to deep neural denoisers. We show, both theoretically and experimentally, that this locality arises directly from pixel correlations present in the image datasets. Moreover, locality patterns are drastically different on specialized datasets, approximating principal components of the data's covariance. We use these insights to craft an analytical denoiser that better matches scores predicted by a deep diffusion model than prior expert-crafted alternatives.

Our key takeaway is that while neural network architectures influence generation quality, their primary role is to capture locality patterns inherent in the data.

Recent work investigates this paradox and proposes changes to the optimal denoiser to close the gap between theory and practice [12,19,25,26]. Kamb and Ganguli [12] hypothesize that inductive biases of the neural network architecture-particularly shift equivariance and locality biases of 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

We discuss preliminaries and related work for our primary line of inquiry, building analytical models of deep diffusion networks.

Denoising diffusion models. Score-based image generative models [8,27,28] learn to reverse the process of adding Gaussian noise to clean data. During training, we sample a data point x 0 from the training data distribution X, a noise level t from the interval [0, 1], and a Gaussian noise direction ϵ ∼ N (0, I); a noise schedule α t is chosen such that α 0 = 1 and α 1 = 0. We then add noise to x 0 to obtain x t = √ α t x 0 + √ 1 -α t ϵ. The training objective for an image diffusion model f (x, t), also

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper11

问问它们各自怎么用它

它引用的顶会 Paper13

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖