An analytic theory of creativity in convolutional diffusion models
Mason Kamb, Surya Ganguli
Abstract
We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identify two simple inductive biases, locality and equivariance, that: (1) induce a form of combinatorial creativity by preventing optimal score-matching; (2) result in fully analytic, completely mechanistically interpretable, local score (LS) and equivariant local score (ELS) machines that, (3) after calibrating a single time-dependent hyperparameter can quantitatively predict the outputs of trained convolution only diffusion models (like ResNets and UNets) with high accuracy (median r 2 of 0.95, 0.94, 0.94, 0.96 for our top model on CIFAR10, FashionMNIST, MNIST, and CelebA). Our model reveals a locally consistent patch mosaic mechanism of creativity, in which diffusion models create exponentially many novel images by mixing and matching different local training set patches at different scales and image locations. Our theory also partially predicts the outputs of pre-trained self-attention enabled UNets (median r 2 ∼ 0.77 on CIFAR10), revealing an intriguing role for attention in carving out semantic coherence from local patch mosaics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers40
- Fast and Scalable Analytical DiffusionXinyi Shang, Peng Sun, Jingyu Lin, Zhiqiang ShenICML 2026 · 1,092 citations
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingTony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc MézardNeurIPS 2025 · 93 citations
- On the Closed-Form of Flow Matching: Generalization Does Not Arise from Target StochasticityQuentin Bertrand, Anne Gagneux, Mathurin Massias, Rémi EmonetNeurIPS 2025 · 49 citations
- Locality in Image Diffusion Models Emerges from Data StatisticsArtem Lukoianov, Chenyang Yuan, Justin M. Solomon, Vincent SitzmannNeurIPS 2025 · 32 citations
- An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion ModelsBinxu Wang, Cengiz PehlevanNeurIPS 2025 · 26 citations
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- On Inductive Biases That Enable Generalization in Diffusion TransformersJie An, De Wang, Pengsheng Guo, Jiebo Luo et al.NeurIPS 2025 · 1 citation
- Towards a Mechanistic Explanation of Diffusion Model GeneralizationMatthew Niedoba, Berend Zwartsenberg, Kevin Patrick Murphy, Frank WoodICML 2025
- Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into ConvolutionsZiYi Dong, Chengxing Zhou, Weijian Deng, Pengxu Wei et al.ICCV 2025
- Generalization in diffusion models arises from geometry-adaptive harmonic representationsZahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, Stéphane MallatICLR 2024 · 168 citations
- What's the score? Automated Denoising Score Matching for Nonlinear DiffusionsRaghav Singhal, Mark Goldstein, Rajesh RanganathICML 2024 · 8 citations
