Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization
Qing Guo, Junya Chen, Dong Wang, Yuewei Yang, Xinwei Deng, Jing Huang, Lawrence Carin, Fan Li, Chenyang Tao
Abstract
Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning. While featuring superior stability, these estimators crucially depend on costly large-batch training, and they sacrifice bound tightness for variance reduction. To overcome these limitations, we revisit the mathematics of popular variational MI bounds from the lens of unnormalized statistical modeling and convex optimization. Our investigation yields a new unified theoretical framework encompassing popular variational MI bounds, and leads to a new simple and powerful contrastive MI estimator we name FLO. Theoretically, we show that the FLO estimator is tight, and it converges under stochastic gradient descent. Empirically, the FLO estimator overcomes the limitations of its predecessors and learns more efficiently. The utility of FLO is verified using extensive benchmarks, and we further inspire the community with novel applications in meta-learning. Our presentation underscores the foundational importance of variational MI estimation in data-efficient learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Factorized Contrastive Learning: Going Beyond Multi-view RedundancyPaul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou et al.NeurIPS 2023 · 137 citations
- Understanding Contrastive Learning via Distributionally Robust OptimizationJunkang Wu, Jiawei Chen, Jiancan Wu, Wentao Shi et al.NeurIPS 2023 · 55 citations
- On the Surrogate Gap between Contrastive and Supervised LossesHan Bao, Yoshihiro Nagano, Kento NozawaICML 2022 · 27 citations
- Max-Sliced Mutual InformationDor Tsur, Ziv Goldfeld, Kristjan H. GreenewaldNeurIPS 2023 · 20 citations
- Contrastive Representation for Data Filtering in Cross-Domain Offline Reinforcement LearningXiaoyu Wen, Chenjia Bai, Kang Xu, Xudong Yu et al.ICML 2024 · 13 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
Related papers
- Improving Mutual Information Estimation with Annealed and Energy-Based BoundsRob Brekelmans, Sicong Huang, Marzyeh Ghassemi, Greg Ver Steeg et al.ICLR 2022 · 16 citations
- Contrastive Predictive Coding Done Right for Mutual Information EstimationJongha Ryu, Pavan Yeddanapudi, Xiangxiang Xu, Gregory W. WornellICLR 2026 · 1 citation
- Flow-based Variational Mutual Information: Fast and Flexible ApproximationsCaleb Dahlke, Jason PachecoICLR 2025
- Gaussian Mutual Information Maximization for Efficient Graph Self-Supervised Learning: Bridging Contrastive-based to Decorrelation-basedJinyong WenACM MM 2024 · 3 citations
- Rethinking Negative Pairs in Code SearchHaochen Li, Xin Zhou, Anh Tuan Luu, Chunyan MiaoEMNLP 2023 · 6 citations
