Lune

NeurIPS2025顶会

Neural Entropy

Akhil Premkumar

2025年份
9被引次数
3顶会引用

摘要

We explore the connection between deep learning and information theory through the paradigm of diffusion models. A diffusion model converts noise into structured data by reinstating, imperfectly, information that is erased when data was diffused to noise. This information is stored in a neural network during training. We quantify this information by introducing a measure called neural entropy, which is related to the total entropy produced by diffusion. Neural entropy is a function of not just the data distribution, but also the diffusive process itself. Measurements of neural entropy on a few simple image diffusion models reveal that they are extremely efficient at compressing large ensembles of structured data.

How much information is stored in a neural network? As a simple example, consider training a neural network to store an 8-bit grayscale image of dimension H × W pixels. The network learns a smooth map from pixel co-ordinates to grayscale intensity values from H × W bytes of raw data. This is not the total number of bytes of the parameters that constitute the network, and not every image of size H × W contains the same amount of information. But it is reasonable to expect that if we push images of higher and higher resolutions/detail onto the same network, at some point the network will not be able to reproduce the images faithfully The question is even more pertinent in the context of generative models. These models are capable of producing seemingly endless variations of the original training data, say images, but that does not mean the neural network has stored an infinite number of images. Rather, generative models store a distribution of images, call it p d , and the generated samples are points that interpolate the training data in p d . This is similar to how the network from the prior example blends the grayscale intensities between neighboring pixels. So the analogous question to ask is this: how many bytes of data is p d worth? The primary goal of this paper is to answer this question in the context of diffusion-based generative models (hint: it is not simply the Shannon entropy of p d , see App. C.2).

Diffusion models serve as a natural bridge between information theory and machine learning, having been inspired by ideas from non-equilibrium thermodynamics [1], which itself can be viewed as an application of information-theoretic principles to physical systems [2][3][4]. Very briefly, samples from a training dataset are incrementally noised till they are distributed as a generic Gaussian, call it p eq , while a neural network learns to reverse these noising steps. Once trained, the network can transform a random Gaussian vector into a highly structured output that resembles a typical member of the training data. In the continuum limit, the noising and denoising stages become diffusive processes [5,6], the thermodynamic properties of which are well established [7][8][9].

Diffusion gradually wipes out information from p d over time (cf. Fig. 6). The information loss is quantified by the total entropy produced during the process, S tot . Within this framework, we can 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖