Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
Abstract
Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions. Existing methods manifest limitations in leveraging nonverbal information for discerning complex semantics in unsupervised scenarios. This paper introduces a novel unsupervised multimodal clustering method (UMC), making a pioneering contribution to this field. UMC introduces a unique approach to constructing augmentation views for multimodal data, which are then used to perform pre-training to establish well-initialized representations for subsequent clustering. An innovative strategy is proposed to dynamically select high-quality samples as guidance for representation learning, gauged by the density of each sample's nearest neighbors. Besides, it is equipped to automatically determine the optimal value for the top-K parameter in each cluster to refine sample selection. Finally, both high-and low-quality samples are used to learn representations conducive to effective clustering. We build baselines on benchmark multimodal intent and dialogue act datasets. UMC shows remarkable improvements of 2-6% scores in clustering metrics over state-of-the-art methods, marking the first successful endeavor in this domain. The complete code and data are available at https://github.com/thuiar/UMC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly DataLiang Zhao, Shubin Ma, Bo Xu, Qingchen ZhangACM MM 2025 · 1 citation
- Unsupervised Semantic Discovery via Global and Local Semantic Alignment in Multimodal ClusteringZhengzhong Zhu, Pei Zhou, Weihong Du, Shiquan Min et al.AAAI 2026
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
Related papers
- New Intent Discovery with Pre-training and Contrastive LearningYuwei Zhang, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu et al.ACL 2022 · 55 citations
- Discovering New Intents with Deep Aligned ClusteringHanlei Zhang, Hua Xu, Ting-En Lin, Rui LyuAAAI 2021 · 138 citations
- Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification with K-Means FeaturesTao Gui, Lizhi Qing, Qi Zhang, Jiacheng Ye et al.AAAI 2020 · 2 citations
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi et al.AAAI 2026
- Deep Multiview Clustering by Contrasting Cluster AssignmentsJie Chen, Hua Mao, Wai Lok Woo, Xi PengICCV 2023 · 142 citations
