Locality in Image Diffusion Models Emerges from Data Statistics
Artem Lukoianov, Chenyang Yuan, Justin M. Solomon, Vincent Sitzmann
Abstract
Recent work has shown that the generalization ability of image diffusion models arises from the locality properties of the trained neural network. In particular, when denoising a particular pixel, the model relies on a limited neighborhood of the input image around that pixel, which, according to the previous work, is tightly related to the ability of these models to produce novel images. Since locality is central to generalization, it is crucial to understand why diffusion models learn local behavior in the first place, as well as the factors that govern the properties of locality patterns. In this work, we present evidence that the locality in deep diffusion models emerges as a statistical property of the image dataset and is not due to the inductive bias of convolutional neural networks, as suggested in previous work. Specifically, we demonstrate that an optimal parametric linear denoiser exhibits similar locality properties to deep neural denoisers. We show, both theoretically and experimentally, that this locality arises directly from pixel correlations present in the image datasets. Moreover, locality patterns are drastically different on specialized datasets, approximating principal components of the data's covariance. We use these insights to craft an analytical denoiser that better matches scores predicted by a deep diffusion model than prior expert-crafted alternatives.
Our key takeaway is that while neural network architectures influence generation quality, their primary role is to capture locality patterns inherent in the data.
Recent work investigates this paradox and proposes changes to the optimal denoiser to close the gap between theory and practice [12,19,25,26]. Kamb and Ganguli [12] hypothesize that inductive biases of the neural network architecture-particularly shift equivariance and locality biases of 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
We discuss preliminaries and related work for our primary line of inquiry, building analytical models of deep diffusion networks.
Denoising diffusion models. Score-based image generative models [8,27,28] learn to reverse the process of adding Gaussian noise to clean data. During training, we sample a data point x 0 from the training data distribution X, a noise level t from the interval [0, 1], and a Gaussian noise direction ϵ ∼ N (0, I); a noise schedule α t is chosen such that α 0 = 1 and α 1 = 0. We then add noise to x 0 to obtain x t = √ α t x 0 + √ 1 -α t ϵ. The training objective for an image diffusion model f (x, t), also
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f42cbbcd-5601-4638-9a30-396456dd34c0Cited by top-tier papers11
- Fast and Scalable Analytical DiffusionXinyi Shang, Peng Sun, Jingyu Lin, Zhiqiang ShenICML 2026 · 1,092 citations
- On the Closed-Form of Flow Matching: Generalization Does Not Arise from Target StochasticityQuentin Bertrand, Anne Gagneux, Mathurin Massias, Rémi EmonetNeurIPS 2025 · 49 citations
- Understanding Representation Dynamics of Diffusion Models via Low-Dimensional ModelingXiao Li, Zekai Zhang, Xiang Li, Siyi Chen et al.NeurIPS 2025 · 19 citations
- Provable Separations between Memorization and Generalization in Diffusion ModelsZeqi Ye, Qijie Zhu, Molei Tao, Minshuo ChenICLR 2026 · 15 citations
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi et al.ICLR 2026 · 14 citations
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Generalization in diffusion models arises from geometry-adaptive harmonic representationsZahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, Stéphane MallatICLR 2024 · 168 citations
Related papers
- Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian StructureXiang Li, Yixiang Dai, Qing QuNeurIPS 2024 · 45 citations
- Towards a Mechanistic Explanation of Diffusion Model GeneralizationMatthew Niedoba, Berend Zwartsenberg, Kevin Patrick Murphy, Frank WoodICML 2025
- On Inductive Biases That Enable Generalization in Diffusion TransformersJie An, De Wang, Pengsheng Guo, Jiebo Luo et al.NeurIPS 2025 · 1 citation
- An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion ModelsBinxu Wang, Cengiz PehlevanNeurIPS 2025 · 26 citations
- An analytic theory of creativity in convolutional diffusion modelsMason Kamb, Surya GanguliICML 2025
