Lune

NeurIPS2025Top-tier venue

Neural Entropy

Akhil Premkumar

2025Year
9Citations
3Top-tier citations

Abstract

We explore the connection between deep learning and information theory through the paradigm of diffusion models. A diffusion model converts noise into structured data by reinstating, imperfectly, information that is erased when data was diffused to noise. This information is stored in a neural network during training. We quantify this information by introducing a measure called neural entropy, which is related to the total entropy produced by diffusion. Neural entropy is a function of not just the data distribution, but also the diffusive process itself. Measurements of neural entropy on a few simple image diffusion models reveal that they are extremely efficient at compressing large ensembles of structured data.

How much information is stored in a neural network? As a simple example, consider training a neural network to store an 8-bit grayscale image of dimension H × W pixels. The network learns a smooth map from pixel co-ordinates to grayscale intensity values from H × W bytes of raw data. This is not the total number of bytes of the parameters that constitute the network, and not every image of size H × W contains the same amount of information. But it is reasonable to expect that if we push images of higher and higher resolutions/detail onto the same network, at some point the network will not be able to reproduce the images faithfully The question is even more pertinent in the context of generative models. These models are capable of producing seemingly endless variations of the original training data, say images, but that does not mean the neural network has stored an infinite number of images. Rather, generative models store a distribution of images, call it p d , and the generated samples are points that interpolate the training data in p d . This is similar to how the network from the prior example blends the grayscale intensities between neighboring pixels. So the analogous question to ask is this: how many bytes of data is p d worth? The primary goal of this paper is to answer this question in the context of diffusion-based generative models (hint: it is not simply the Shannon entropy of p d , see App. C.2).

Diffusion models serve as a natural bridge between information theory and machine learning, having been inspired by ideas from non-equilibrium thermodynamics [1], which itself can be viewed as an application of information-theoretic principles to physical systems [2][3][4]. Very briefly, samples from a training dataset are incrementally noised till they are distributed as a generic Gaussian, call it p eq , while a neural network learns to reverse these noising steps. Once trained, the network can transform a random Gaussian vector into a highly structured output that resembles a typical member of the training data. In the continuum limit, the noising and denoising stages become diffusive processes [5,6], the thermodynamic properties of which are well established [7][8][9].

Diffusion gradually wipes out information from p d over time (cf. Fig. 6). The information loss is quantified by the total entropy produced during the process, S tot . Within this framework, we can 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 67b950fe-b4cf-49a1-8896-4077d0bc8c0c

Cited by top-tier papers3

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines