Learning Optimal Multimodal Information Bottleneck Representations
Qilong Wu, Yiyang Shao, Jun Wang, Xiaobo Sun
摘要
Leveraging high-quality joint representations from multimodal data can greatly enhance model performance in various machine-learning based applications. Recent multimodal learning methods, based on the multimodal information bottleneck (MIB) principle, aim to generate optimal MIB with maximal task-relevant information and minimal superfluous information via regularization. However, these methods often set ad hoc regularization weights and overlook imbalanced task-relevant information across modalities, limiting their ability to achieve optimal MIB. To address this gap, we propose a novel multimodal learning framework, Optimal Multimodal Information Bottleneck (OMIB), whose optimization objective guarantees the achievability of optimal MIB by setting the regularization weight within a theoretically derived bound. OMIB further addresses imbalanced task-relevant information by dynamically adjusting regularization weights per modality, promoting the inclusion of all task-relevant information. Moreover, we establish a solid information-theoretical foundation for OMIB's optimization and implement it under the variational approximation framework for computational efficiency. Finally, we empirically validate the OMIB's theoretical properties on synthetic data and demonstrate its superiority over the state-of-the-art benchmark methods in various downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Comprehensive Information-Decomposition Analysis of Large Vision-Language ModelsLixin Xiu, Xufang Luo, Hideki NakayamaICLR 2026 · 被引用 4 次
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva 等ICML 2026
- Structured Multi-modal Graph Disentanglement for Psychiatric DiagnosisHongyu Shi, Kaizhong Zheng, WS Zhai, Shuai Jiang 等ICML 2026
它引用的顶会 Paper4
- Multi-View Information-Bottleneck Representation LearningZhibin Wan, Changqing Zhang, Pengfei Zhu, Qinghua HuAAAI 2021 · 被引用 116 次
- The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge DistillationZihui Xue, Zhengqi Gao, Sucheng Ren, Hang ZhaoICLR 2023 · 被引用 12 次
- Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-IdentificationXudong Tian, Zhizhong Zhang, Shaohui Lin, Yanyun Qu 等CVPR 2021
- An Information Criterion for Controlled Disentanglement of Multimodal DataChenyu Wang, Sharut Gupta, Xinyi Zhang, Sana Tonekaboni 等ICLR 2025
相关 Paper
- Aligning Multimodal Representations through an Information BottleneckAntonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer 等ICML 2025
- RedCore: Relative Advantage Aware Cross-Modal Representation Learning for Missing Modalities with Imbalanced Missing RatesJun Sun, Xinxin Zhang, Shoukang Han, Yu-Ping Ruan 等AAAI 2024
- Multi-aspect Self-guided Deep Information Bottleneck for Multi-modal ClusteringShizhe Hu, Jiahao Fan, Guoliang Zou, Yangdong YeAAAI 2025 · 被引用 5 次
- Exploring the Trade-Off within Visual Information for MultiModal Sentence SummarizationMinghuan Yuan, Shiyao Cui, Xinghua Zhang, Shicheng Wang 等SIGIR 2024 · 被引用 3 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
