Improved Mutual Information Estimation
Youssef Mroueh, Igor Melnyk, Pierre L. Dognin, Jarret Ross, Tom Sercu
Abstract
We propose a new variational lower bound on the KL divergence and show that the Mutual Information (MI) can be estimated by maximizing this bound using a witness function on a hypothesis function class and an auxiliary scalar variable. If the function class is in a Reproducing Kernel Hilbert Space (RKHS), this leads to a jointly convex problem. We analyze the bound by deriving its dual formulation and show its connection to a likelihood ratio estimation problem. We show that the auxiliary variable introduced in our variational form plays the role of a Lagrange multiplier that enforces a normalization constraint on the likelihood ratio. By extending the function space to neural networks, we propose an efficient neural MI estimator, and validate its performance on synthetic examples, showing advantage over the existing baselines. We then demonstrate the strength of our estimator in large-scale self-supervised representation learning through MI maximization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fa3fe8b-d1b3-4e7a-9adf-5cd7efc06196Cited by top-tier papers4
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong et al.ICCV 2023 · 57 citations
- Dual Projection Generative Adversarial Networks for Conditional Image GenerationLigong Han, Martin Renqiang Min, Anastasis Stathopoulos, Yu Tian et al.ICCV 2021 · 22 citations
- Information Bottleneck Analysis of Deep Neural Networks via Lossy CompressionIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya et al.ICLR 2024 · 20 citations
- Reducing Sentiment Bias in Pre-trained Sentiment Classification via Adaptive Gumbel AttackJiachen Tian, Shizhan Chen, Xiaowang Zhang, Xin Wang et al.AAAI 2023 · 5 citations
Builds on1
Related papers
- Reliable Estimation of KL Divergence using a Discriminator in Reproducing Kernel Hilbert SpaceSandesh Ghimire, Aria Masoomi, Jennifer G. DyNeurIPS 2021 · 17 citations
- Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation LearningReuben Dorent, Polina Golland, William (Sandy) WellsNeurIPS 2025 · 7 citations
- Gaussian Mutual Information Maximization for Efficient Graph Self-Supervised Learning: Bridging Contrastive-based to Decorrelation-basedJinyong WenACM MM 2024 · 3 citations
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 243 citations
- Generative Particle Variational Inference via Estimation of Functional GradientsNeale Ratzlaff, Qinxun Bai, Fuxin Li, Wei XuICML 2021
