Deep Mutual Information Maximin for Cross-Modal Clustering
Yiqiao Mao, Xiaoqiang Yan, Qiang Guo, Yangdong Ye
Abstract
Cross-modal clustering (CMC) aims to enhance the clustering performance by exploring complementary information from multiple modalities. However, the performances of existing CMC algorithms are still unsatisfactory due to the conflict of heterogeneous modalities and the high-dimensional non-linear property of individual modality. In this paper, a novel deep mutual information maximin (DMIM) method for cross-modal clustering is proposed to maximally preserve the shared information of multiple modalities while eliminating the superfluous information of individual modalities in an end-to-end manner. Specifically, a multi-modal shared encoder is firstly built to align the latent feature distributions by sharing parameters across modalities. Then, DMIM formulates the complementarity of multi-modalities representations as an mutual information maximin objective function, in which the shared information of multiple modalities and the superfluous information of individual modalities are identified by mutual information maximization and minimization respectively. To solve the DMIM objective function, we propose a variational optimization method to ensure it converge to a local optimal solution. Moreover, an auxiliary overclustering mechanism is employed to optimize the clustering structure by introducing more detailed clustering classes. Extensive experimental results demonstrate the superiority of DMIM method over the state-of-the-art cross-modal clustering methods on IAPR-TC12, ESP-Game, MIRFlickr and NUS-Wide datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 025524c9-ddb1-45f3-a80f-32d3a8d5e062Cited by top-tier papers15
- Differentiable Information Bottleneck for Deterministic Multi-View ClusteringXiaoqiang Yan, Zhixiang Jin, Fengshou Han, Yangdong YeCVPR 2024 · 19 citations
- Few-shot Continual Infomax LearningZiqi Gu, Chunyan Xu, Jian Yang, Zhen CuiICCV 2023 · 18 citations
- Self-supervised Trusted Contrastive Multi-view Clustering with Uncertainty RefinedShizhe Hu, Binyan Tian, Weibo Liu, Yangdong YeAAAI 2025 · 11 citations
- Live and Learn: Continual Action Clustering with Incremental ViewsXiaoqiang Yan, Yingtao Gan, Yiqiao Mao, Yangdong Ye et al.AAAI 2024 · 11 citations
- Multi-Level Cross-Modal Alignment for Image ClusteringLiping Qiu, Qin Zhang, Xiaojun Chen, Shaotian CaiAAAI 2024 · 8 citations
Builds on7
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
- Multi-View Clustering in Latent Embedding SpaceMan-Sheng Chen, Ling Huang, Chang-Dong Wang, Dong HuangAAAI 2020 · 275 citations
- Tensor-SVD Based Graph Learning for Multi-View Subspace ClusteringQuanxue Gao, Wei Xia, Zhizhen Wan, De-Yan Xie et al.AAAI 2020 · 231 citations
Related papers
- Diversity-oriented Deep Multi-modal ClusteringYanzheng Wang, Xin Yang, Yujun Wang, Shizhe Hu et al.NeurIPS 2025
- Multi-aspect Self-guided Deep Information Bottleneck for Multi-modal ClusteringShizhe Hu, Jiahao Fan, Guoliang Zou, Yangdong YeAAAI 2025 · 5 citations
- Super Deep Contrastive Information Bottleneck for Multi-modal ClusteringZhengzheng Lou, Ke Zhang, Yucong Wu, Shizhe HuICML 2025
- End-to-End Adversarial-Attention Network for Multi-Modal ClusteringRunwu Zhou, Yi-Dong ShenCVPR 2020
- Geometry-Aware Variational Information Maximization for Deep Incomplete Multi-view ClusteringWenlan Chen, Lu Gao, Daoyuan Wang, Fei Guo et al.AAAI 2026
