Super Deep Contrastive Information Bottleneck for Multi-modal Clustering
Zhengzheng Lou, Ke Zhang, Yucong Wu, Shizhe Hu
Abstract
In an era of increasingly diverse information sources, multi-modal clustering (MMC) has become a key technology for processing multimodal data. It can apply and integrate the feature information and potential relationships of different modalities. Although there is a wealth of research on MMC, due to the complexity of datasets, a major challenge remains in how to deeply explore the complex latent information and interdependencies between modalities. To address this issue, this paper proposes a method called super deep contrastive information bottleneck (SDCIB) for MMC, which aims to explore and utilize all types of latent information to the fullest extent. Specifically, the proposed SDCIB explicitly introduces the rich information contained in the encoder's hidden layers into the loss function for the first time, thoroughly mining both modal features and the hidden relationships between modalities. Moreover, the proposed SDCIB performs dual optimization by simultaneously considering consistency information from both the feature distribution and clustering assignment perspectives, the proposed SDCIB significantly improves clustering accuracy and robustness. We conducted experiments on 4 multi-modal datasets and the accuracy of the method on the ESP dataset improved by 9.3%. The results demonstrate the superiority and clever design of the proposed SDCIB. The source code is available on https://github.com/ShizheHu .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Calibrated Information Bottleneck for Trusted Multi-modal ClusteringShizhe Hu, Zhangwen Gou, Shuaiju Li, Jin Qin et al.ICLR 2026
- Information-Theoretic Disentangled Latent Modeling with Conditional Diffusion for Incomplete Multi-View ClusteringWenlan Chen, Lu Gao, Daoyuan Wang, Cheng Liang et al.ICML 2026
- Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View ClusteringTaichun Zhou, Zhibin Dong, Hao Tan, Siwei Wang et al.CVPR 2026
Builds on13
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Multi-level Feature Learning for Contrastive Multi-view ClusteringJie Xu, Huayi Tang, Yazhou Ren, Liang Peng et al.CVPR 2022 · 335 citations
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
- DealMVC: Dual Contrastive Calibration for Multi-view ClusteringXihong Yang, Jiaqi Jin, Siwei Wang, Ke Liang et al.ACM MM 2023 · 138 citations
- Incomplete Contrastive Multi-View Clustering with High-Confidence GuidingGuoqing Chao, Yi Jiang, Dianhui ChuAAAI 2024 · 135 citations
Related papers
- Multi-aspect Self-guided Deep Information Bottleneck for Multi-modal ClusteringShizhe Hu, Jiahao Fan, Guoliang Zou, Yangdong YeAAAI 2025 · 5 citations
- Deep Mutual Information Maximin for Cross-Modal ClusteringYiqiao Mao, Xiaoqiang Yan, Qiang Guo, Yangdong YeAAAI 2021 · 58 citations
- Diversity-oriented Deep Multi-modal ClusteringYanzheng Wang, Xin Yang, Yujun Wang, Shizhe Hu et al.NeurIPS 2025
- End-to-End Adversarial-Attention Network for Multi-Modal ClusteringRunwu Zhou, Yi-Dong ShenCVPR 2020
- Cross-Modal Subspace Clustering via Deep Canonical Correlation AnalysisQuanxue Gao, Huanhuan Lian, Qianqian Wang, Gan SunAAAI 2020 · 65 citations
