Super Deep Contrastive Information Bottleneck for Multi-modal Clustering
Zhengzheng Lou, Ke Zhang, Yucong Wu, Shizhe Hu
摘要
In an era of increasingly diverse information sources, multi-modal clustering (MMC) has become a key technology for processing multimodal data. It can apply and integrate the feature information and potential relationships of different modalities. Although there is a wealth of research on MMC, due to the complexity of datasets, a major challenge remains in how to deeply explore the complex latent information and interdependencies between modalities. To address this issue, this paper proposes a method called super deep contrastive information bottleneck (SDCIB) for MMC, which aims to explore and utilize all types of latent information to the fullest extent. Specifically, the proposed SDCIB explicitly introduces the rich information contained in the encoder's hidden layers into the loss function for the first time, thoroughly mining both modal features and the hidden relationships between modalities. Moreover, the proposed SDCIB performs dual optimization by simultaneously considering consistency information from both the feature distribution and clustering assignment perspectives, the proposed SDCIB significantly improves clustering accuracy and robustness. We conducted experiments on 4 multi-modal datasets and the accuracy of the method on the ESP dataset improved by 9.3%. The results demonstrate the superiority and clever design of the proposed SDCIB. The source code is available on https://github.com/ShizheHu .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Calibrated Information Bottleneck for Trusted Multi-modal ClusteringShizhe Hu, Zhangwen Gou, Shuaiju Li, Jin Qin 等ICLR 2026
- Information-Theoretic Disentangled Latent Modeling with Conditional Diffusion for Incomplete Multi-View ClusteringWenlan Chen, Lu Gao, Daoyuan Wang, Cheng Liang 等ICML 2026
- Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View ClusteringTaichun Zhou, Zhibin Dong, Hao Tan, Siwei Wang 等CVPR 2026
它引用的顶会 Paper13
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Multi-level Feature Learning for Contrastive Multi-view ClusteringJie Xu, Huayi Tang, Yazhou Ren, Liang Peng 等CVPR 2022 · 被引用 335 次
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- DealMVC: Dual Contrastive Calibration for Multi-view ClusteringXihong Yang, Jiaqi Jin, Siwei Wang, Ke Liang 等ACM MM 2023 · 被引用 138 次
- Incomplete Contrastive Multi-View Clustering with High-Confidence GuidingGuoqing Chao, Yi Jiang, Dianhui ChuAAAI 2024 · 被引用 135 次
相关 Paper
- Multi-aspect Self-guided Deep Information Bottleneck for Multi-modal ClusteringShizhe Hu, Jiahao Fan, Guoliang Zou, Yangdong YeAAAI 2025 · 被引用 5 次
- Deep Mutual Information Maximin for Cross-Modal ClusteringYiqiao Mao, Xiaoqiang Yan, Qiang Guo, Yangdong YeAAAI 2021 · 被引用 58 次
- Diversity-oriented Deep Multi-modal ClusteringYanzheng Wang, Xin Yang, Yujun Wang, Shizhe Hu 等NeurIPS 2025
- End-to-End Adversarial-Attention Network for Multi-Modal ClusteringRunwu Zhou, Yi-Dong ShenCVPR 2020
- Cross-Modal Subspace Clustering via Deep Canonical Correlation AnalysisQuanxue Gao, Huanhuan Lian, Qianqian Wang, Gan SunAAAI 2020 · 被引用 65 次
