Neural Entropy
Akhil Premkumar
Abstract
We explore the connection between deep learning and information theory through the paradigm of diffusion models. A diffusion model converts noise into structured data by reinstating, imperfectly, information that is erased when data was diffused to noise. This information is stored in a neural network during training. We quantify this information by introducing a measure called neural entropy, which is related to the total entropy produced by diffusion. Neural entropy is a function of not just the data distribution, but also the diffusive process itself. Measurements of neural entropy on a few simple image diffusion models reveal that they are extremely efficient at compressing large ensembles of structured data.
How much information is stored in a neural network? As a simple example, consider training a neural network to store an 8-bit grayscale image of dimension H × W pixels. The network learns a smooth map from pixel co-ordinates to grayscale intensity values from H × W bytes of raw data. This is not the total number of bytes of the parameters that constitute the network, and not every image of size H × W contains the same amount of information. But it is reasonable to expect that if we push images of higher and higher resolutions/detail onto the same network, at some point the network will not be able to reproduce the images faithfully The question is even more pertinent in the context of generative models. These models are capable of producing seemingly endless variations of the original training data, say images, but that does not mean the neural network has stored an infinite number of images. Rather, generative models store a distribution of images, call it p d , and the generated samples are points that interpolate the training data in p d . This is similar to how the network from the prior example blends the grayscale intensities between neighboring pixels. So the analogous question to ask is this: how many bytes of data is p d worth? The primary goal of this paper is to answer this question in the context of diffusion-based generative models (hint: it is not simply the Shannon entropy of p d , see App. C.2).
Diffusion models serve as a natural bridge between information theory and machine learning, having been inspired by ideas from non-equilibrium thermodynamics [1], which itself can be viewed as an application of information-theoretic principles to physical systems [2][3][4]. Very briefly, samples from a training dataset are incrementally noised till they are distributed as a generic Gaussian, call it p eq , while a neural network learns to reverse these noising steps. Once trained, the network can transform a random Gaussian vector into a highly structured output that resembles a typical member of the training data. In the continuum limit, the noising and denoising stages become diffusive processes [5,6], the thermodynamic properties of which are well established [7][8][9].
Diffusion gradually wipes out information from p d over time (cf. Fig. 6). The information loss is quantified by the total entropy produced during the process, S tot . Within this framework, we can 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67b950fe-b4cf-49a1-8896-4077d0bc8c0cCited by top-tier papers3
- Entropic Time Schedulers for Generative Diffusion ModelsDejan Stancevic, Florian Handke, Luca AmbrogioniNeurIPS 2025 · 19 citations
- The Entropic Signature of Class Speciation in Diffusion ModelsFlorian Handke, Dejan Stancevic, Felix Koulischer, Thomas Demeester et al.ICML 2026 · 5 citations
- On the Separability of Information in Diffusion ModelsAkhil PremkumarICML 2026 · 1 citation
Builds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Quantification and Analysis of Layer-wise and Pixel-wise Information DiscardingHaotian Ma, Hao Zhang, Fan Zhou, Yinqing Zhang et al.ICML 2022 · 2 citations
- An Information-Theoretic Regularizer for Lossy Neural Image CompressionYingwen Zhang, Meng Wang, Xihua Sheng, Peilin Chen et al.ICCV 2025
- Improved Sample Complexity Bounds for Diffusion Model TrainingShivam Gupta, Aditya Parulekar, Eric Price, Zhiyang XunNeurIPS 2024 · 16 citations
- Information Bottleneck Analysis of Deep Neural Networks via Lossy CompressionIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya et al.ICLR 2024 · 20 citations
- Local Intrinsic Dimensional EntropyRohan Ghosh, Mehul MotaniAAAI 2023 · 2 citations
