Improving Mutual Information Estimation with Annealed and Energy-Based Bounds
Rob Brekelmans, Sicong Huang, Marzyeh Ghassemi, Greg Ver Steeg, Roger Baker Grosse, Alireza Makhzani
摘要
Mutual information (MI) is a fundamental quantity in information theory and machine learning. However, direct estimation of MI is intractable, even if the true joint probability density for the variables of interest is known, as it involves estimating a potentially high-dimensional log partition function. In this work, we present a unifying view of existing MI bounds from the perspective of importance sampling, and propose three novel bounds based on this approach. Since accurate estimation of MI without density information requires a sample size exponential in the true MI, we assume either a single marginal or the full joint density information is known. In settings where the full joint density is available, we propose Multi-Sample Annealed Importance Sampling (AIS) bounds on MI, which we demonstrate can tightly estimate large values of MI in our experiments. In settings where only a single marginal distribution is known, we propose Generalized IWAE (GIWAE) and MINE-AIS bounds. Our GIWAE bound unifies variational and contrastive bounds in a single framework that generalizes InfoNCE, IWAE, and Barber-Agakov bounds. Our MINE-AIS method improves upon existing energy-based methods such as MINE-DV and MINE-F by directly optimizing a tighter lower bound on MI. MINE-AIS uses MCMC sampling to estimate gradients for training and Multi-Sample AIS for evaluating the bound. Our methods are particularly suitable for evaluating MI in deep generative models, since explicit forms of the marginal or joint densities are often available. We evaluate our bounds on estimating the MI of VAEs and GANs trained on the MNIST and CIFAR datasets, and showcase significant gains over existing bounds in these challenging settings with high ground truth MI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Score-Based Diffusion meets Annealed Importance SamplingArnaud Doucet, Will Grathwohl, Alexander G. de G. Matthews, Heiko StrathmannNeurIPS 2022 · 被引用 68 次
- Probabilistic Inference in Language Models via Twisted Sequential Monte CarloStephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Baker GrosseICML 2024 · 被引用 61 次
- Interpretable Diffusion via Information DecompositionXianghao Kong, Ollie Liu, Han Li, Dani Yogatama 等ICLR 2024 · 被引用 37 次
- Tight Mutual Information Estimation With Contrastive Fenchel-Legendre OptimizationQing Guo, Junya Chen, Dong Wang, Yuewei Yang 等NeurIPS 2022 · 被引用 28 次
- MINDE: Mutual Information Neural Diffusion EstimationGiulio Franzese, Mustapha Bounoua, Pietro MichiardiICLR 2024 · 被引用 22 次
它引用的顶会 Paper3
- Generalized Energy Based ModelsMichael Arbel, Liang Zhou, Arthur GrettonICLR 2021 · 被引用 254 次
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 被引用 243 次
- Evaluating Lossy Compression Rates of Deep Generative ModelsSicong Huang, Alireza Makhzani, Yanshuai Cao, Roger B. GrosseICML 2020 · 被引用 30 次
相关 Paper
- Efficient Mixture Learning in Black-Box Variational InferenceAlexandra Hotti, Oskar Kviman, Ricky Molén, Víctor Elvira 等ICML 2024 · 被引用 3 次
- Monte Carlo Variational Auto-EncodersAchille Thin, Nikita Kotelevskii, Arnaud Doucet, Alain Durmus 等ICML 2021 · 被引用 51 次
- Contrastive Predictive Coding Done Right for Mutual Information EstimationJongha Ryu, Pavan Yeddanapudi, Xiangxiang Xu, Gregory W. WornellICLR 2026 · 被引用 1 次
- Flow-based Variational Mutual Information: Fast and Flexible ApproximationsCaleb Dahlke, Jason PachecoICLR 2025
- Mutual Information Gradient Estimation for Representation LearningLiangjian Wen, Yiji Zhou, Lirong He, Mingyuan Zhou 等ICLR 2020 · 被引用 34 次
