Unsupervised Semantic Discovery via Global and Local Semantic Alignment in Multimodal Clustering
Zhengzhong Zhu, Pei Zhou, Weihong Du, Shiquan Min, Jiangping Zhu
Abstract
Unsupervised multimodal semantic discovery aims to learn discriminative representations from multimodal data. However, existing methods suffer from two key limitations. First, they only align instances across modalities without modeling semantic-level consistency, which fails to mitigate semantic bias caused by the gaps among feature distributions of multiple modalities. Second, they inevitably generate incorrect negative pairs during contrastive learning, pushing semantically similar samples apart. To address these challenges, we propose GLAD (Global and Local semantic Alignment for unsupervised multimodal semantic Discovery), which aligns multimodal data at both global and local semantic levels. At the global level, GSA integrates multi-modal features into a shared space and employs joint clustering via optimal transport to capture common semantic patterns while mitigating cross-modality semantic bias. At the local level, LSA adaptively weights samples within each cluster based on their semantic importance, alleviating the effect of incorrect negative pairs. Through the joint optimization of GSA and LSA, GLAD effectively captures both the global semantic structure and the local semantic nuances of multimodal data. Extensive experiments on three benchmark datasets demonstrate GLAD significantly outperforms state-of-the-art methods, with an average improvement of 3.22%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 548890fb-acd1-4c5e-8702-bde0fa30aa4bCited by top-tier papers2
- Fine-to-Coarse Fairness-Informed Multi-View ClusteringShengju Yu, Suyuan Liu, Wenhao SHAO, Siwei Wang et al.ICML 2026
- Homophily-Heterogeneity Gradient Surgery for Federated Graph LearningSujia Huang, Lele Fu, Shunxin Xiao, Xiaoya Zhang et al.ICML 2026
Builds on14
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- Discovering New Intents with Deep Aligned ClusteringHanlei Zhang, Hua Xu, Ting-En Lin, Rui LyuAAAI 2021 · 138 citations
- Discovering New Intents via Constrained Deep Adaptive Clustering with Cluster RefinementTing-En Lin, Hua Xu, Hanlei ZhangAAAI 2020 · 127 citations
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne et al.ICCV 2021 · 98 citations
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
Related papers
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao et al.CVPR 2026
- Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal AlignmentLinlin Ji, Li LiuAAAI 2026
- Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive LearningZhicheng Yang, Yichen Liu, Chang Ge, Xiaopeng JiangCVPR 2026
- AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationHaifeng Zhong, Fan Tang, Zhuo Chen, Hyung Jin Chang et al.ICCV 2025 · 9 citations
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
