Unsupervised Semantic Discovery via Global and Local Semantic Alignment in Multimodal Clustering
Zhengzhong Zhu, Pei Zhou, Weihong Du, Shiquan Min, Jiangping Zhu
摘要
Unsupervised multimodal semantic discovery aims to learn discriminative representations from multimodal data. However, existing methods suffer from two key limitations. First, they only align instances across modalities without modeling semantic-level consistency, which fails to mitigate semantic bias caused by the gaps among feature distributions of multiple modalities. Second, they inevitably generate incorrect negative pairs during contrastive learning, pushing semantically similar samples apart. To address these challenges, we propose GLAD (Global and Local semantic Alignment for unsupervised multimodal semantic Discovery), which aligns multimodal data at both global and local semantic levels. At the global level, GSA integrates multi-modal features into a shared space and employs joint clustering via optimal transport to capture common semantic patterns while mitigating cross-modality semantic bias. At the local level, LSA adaptively weights samples within each cluster based on their semantic importance, alleviating the effect of incorrect negative pairs. Through the joint optimization of GSA and LSA, GLAD effectively captures both the global semantic structure and the local semantic nuances of multimodal data. Extensive experiments on three benchmark datasets demonstrate GLAD significantly outperforms state-of-the-art methods, with an average improvement of 3.22%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fine-to-Coarse Fairness-Informed Multi-View ClusteringShengju Yu, Suyuan Liu, Wenhao SHAO, Siwei Wang 等ICML 2026
- Homophily-Heterogeneity Gradient Surgery for Federated Graph LearningSujia Huang, Lele Fu, Shunxin Xiao, Xiaoya Zhang 等ICML 2026
它引用的顶会 Paper14
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du 等ACM MM 2022 · 被引用 260 次
- Discovering New Intents with Deep Aligned ClusteringHanlei Zhang, Hua Xu, Ting-En Lin, Rui LyuAAAI 2021 · 被引用 138 次
- Discovering New Intents via Constrained Deep Adaptive Clustering with Cluster RefinementTing-En Lin, Hua Xu, Hanlei ZhangAAAI 2020 · 被引用 127 次
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne 等ICCV 2021 · 被引用 98 次
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 等ACM MM 2022 · 被引用 66 次
相关 Paper
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao 等CVPR 2026
- Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal AlignmentLinlin Ji, Li LiuAAAI 2026
- Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive LearningZhicheng Yang, Yichen Liu, Chang Ge, Xiaopeng JiangCVPR 2026
- AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationHaifeng Zhong, Fan Tang, Zhuo Chen, Hyung Jin Chang 等ICCV 2025 · 被引用 9 次
- UDCH: Unsupervised Dynamic Weighted Cluster-cooperative Hashing for Cross-modal RetreivalYuanzhi Zhao, Fan Yang, Yudong Zhao, Xiaoyu LiAAAI 2026
