Lune

ICLR2026Top-tier venue

Information Estimation with Discrete Diffusion

Alberto Foresti, Giulio Franzese, Pietro Michiardi

2026Year

Abstract

Information-theoretic measures, such as Mutual Information (MI), play a crucial role in understanding non-linear relationships between random variables and are widely used across scientific disciplines. Yet, their use on real-world discrete data remains challenging. Existing methods typically rely on embedding discrete data into a continuous space and apply neural estimators originally designed for continuous distributions. This process requires careful engineering for both the embedding model and estimator architecture, but suffers from issues related to high data dimensionality. In this work, we introduce INFO-SEDD, a discrete diffusion-based approach that bridges information-theoretic estimation and generative modeling such that they can be used to compute Kullback-Leibler divergences. Backed by Continuous Time Markov Chains theory principles, the design of INFO-SEDD is lightweight and scalable and allows seamless integration with pretrained models. We showcase the versatility of our approach through applications on motif discovery in genetic promoter data, semantic-aware model selection in text summarization, and entropy estimation in Ising models. Finally, we construct consistency tests on real-world textual and genomics data. Our experiments demonstrate that INFO-SEDD outperforms alternatives that rely on the "embedding trick". Our results position INFO-SEDD as a robust and scalable tool for information-theoretic analysis of discrete data. We provide the open source code at https://github.com/AlbertoForesti/mutinfo-diffusion.

Published as a conference paper at ICLR 2026 posed in the literature. While classical estimators for discrete random variables exist (Pinchas et al., 2024), their accuracy rapidly decreases with increasing data dimensionality. Applications that would benefit from scalable estimators of MI include, among others, DNA or peptide sequencing (Newcomb and Sayood, 2021;Xia et al.), text summarization (Darrin et al., 2024) and neuroscience (Chai et al., 2009), to name a few examples. Consequently, the development of new estimation techniques is of paramount importance for the broader scientific community.

A common workaround to deal with high-dimensional, discrete data is to embed it in a continuous space and use neural estimators conceived for continuous distributions. One recent example is (Lee and Rhee, 2024), where it is shown that the embeddings of pretrained language models can provide meaningful representations to estimate information theoretic quantities in unstructured data. However, such a process may not fully capture the discrete nature of the underlying data and might suffer from several limitations, such as the necessity to consider application-specific embeddings. For example, we show that MINDE (Franzese et al., 2023), which is a strong MI estimator in continuous domains, like image data, struggles with discrete data.

In this work, we build on the growing literature of diffusion-based estimators (Kong et al., 2022;Franzese et al., 2024), and present INFO-SEDD, a novel method for estimating information theoretic quantities of discrete data using Continuous Time Markov Chains (CTMCs) (Lou et al., 2024). These stochastic processes have recently seen a surge in popularity for applications such as generative language modeling (Lou et al., 2024;Sahoo et al., 2024;Nie et al., 2026). Their fundamental working principle is the reversal of a perturbation process which starts with clean data from a given distribution and that converges to uninformative noise. The workhorse of these approaches is the score function, which contains information about the probability distributions associated with the CTMCs at different time instants. Our proposed method, INFO-SEDD, builds upon such mathematical framework, extending it via Dynkin's lemma (Hanson, 2007), and leverages score functions to compute key information-theoretic metrics, such as MI between two random variables, and the entropy of a given distribution. By carefully selecting perturbation processes, our approach requires training only a single parametric model to compute MI across arbitrary subsets of variables. Furthermore, INFO-SEDD seamlessly integrates with pretrained models, without requiring ad-hoc procedures to work with discrete data.

To rigorously evaluate our method, we perform a large number of experiments both on synthetic and real data, and compare INFO-SEDD to state-of-the-art methods. We design a synthetic benchmark with ground truth MI that presents challenges such as high data dimensionality and high MI scenarios. Our results demonstrate that INFO-SEDD is both robust and consistently outperforms existing estimation methods. We then focus on two application domains in which we compare INFO-SEDD to competitors, which all require the "embedding trick" discussed above. First, we tackle the text summarization domain, evaluate the consistency of MI estimators, and study if MI represents a meaningful signal to perform

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 690f6a79-1cac-416a-8d54-5b3c8651ae9d

Builds on23

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines