Deep Mutual Information Maximin for Cross-Modal Clustering
Yiqiao Mao, Xiaoqiang Yan, Qiang Guo, Yangdong Ye
摘要
Cross-modal clustering (CMC) aims to enhance the clustering performance by exploring complementary information from multiple modalities. However, the performances of existing CMC algorithms are still unsatisfactory due to the conflict of heterogeneous modalities and the high-dimensional non-linear property of individual modality. In this paper, a novel deep mutual information maximin (DMIM) method for cross-modal clustering is proposed to maximally preserve the shared information of multiple modalities while eliminating the superfluous information of individual modalities in an end-to-end manner. Specifically, a multi-modal shared encoder is firstly built to align the latent feature distributions by sharing parameters across modalities. Then, DMIM formulates the complementarity of multi-modalities representations as an mutual information maximin objective function, in which the shared information of multiple modalities and the superfluous information of individual modalities are identified by mutual information maximization and minimization respectively. To solve the DMIM objective function, we propose a variational optimization method to ensure it converge to a local optimal solution. Moreover, an auxiliary overclustering mechanism is employed to optimize the clustering structure by introducing more detailed clustering classes. Extensive experimental results demonstrate the superiority of DMIM method over the state-of-the-art cross-modal clustering methods on IAPR-TC12, ESP-Game, MIRFlickr and NUS-Wide datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Differentiable Information Bottleneck for Deterministic Multi-View ClusteringXiaoqiang Yan, Zhixiang Jin, Fengshou Han, Yangdong YeCVPR 2024 · 被引用 19 次
- Few-shot Continual Infomax LearningZiqi Gu, Chunyan Xu, Jian Yang, Zhen CuiICCV 2023 · 被引用 18 次
- Self-supervised Trusted Contrastive Multi-view Clustering with Uncertainty RefinedShizhe Hu, Binyan Tian, Weibo Liu, Yangdong YeAAAI 2025 · 被引用 11 次
- Live and Learn: Continual Action Clustering with Incremental ViewsXiaoqiang Yan, Yingtao Gan, Yiqiao Mao, Yangdong Ye 等AAAI 2024 · 被引用 11 次
- Multi-Level Cross-Modal Alignment for Image ClusteringLiping Qiu, Qin Zhang, Xiaojun Chen, Shaotian CaiAAAI 2024 · 被引用 8 次
它引用的顶会 Paper7
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 被引用 873 次
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- Multi-View Clustering in Latent Embedding SpaceMan-Sheng Chen, Ling Huang, Chang-Dong Wang, Dong HuangAAAI 2020 · 被引用 275 次
- Tensor-SVD Based Graph Learning for Multi-View Subspace ClusteringQuanxue Gao, Wei Xia, Zhizhen Wan, De-Yan Xie 等AAAI 2020 · 被引用 231 次
相关 Paper
- Diversity-oriented Deep Multi-modal ClusteringYanzheng Wang, Xin Yang, Yujun Wang, Shizhe Hu 等NeurIPS 2025
- Multi-aspect Self-guided Deep Information Bottleneck for Multi-modal ClusteringShizhe Hu, Jiahao Fan, Guoliang Zou, Yangdong YeAAAI 2025 · 被引用 5 次
- Super Deep Contrastive Information Bottleneck for Multi-modal ClusteringZhengzheng Lou, Ke Zhang, Yucong Wu, Shizhe HuICML 2025
- End-to-End Adversarial-Attention Network for Multi-Modal ClusteringRunwu Zhou, Yi-Dong ShenCVPR 2020
- Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view ClusteringWenlan Chen, Lu Gao, Daoyuan Wang, Fei Guo 等AAAI 2026
