Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
摘要
Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions. Existing methods manifest limitations in leveraging nonverbal information for discerning complex semantics in unsupervised scenarios. This paper introduces a novel unsupervised multimodal clustering method (UMC), making a pioneering contribution to this field. UMC introduces a unique approach to constructing augmentation views for multimodal data, which are then used to perform pre-training to establish well-initialized representations for subsequent clustering. An innovative strategy is proposed to dynamically select high-quality samples as guidance for representation learning, gauged by the density of each sample's nearest neighbors. Besides, it is equipped to automatically determine the optimal value for the top-K parameter in each cluster to refine sample selection. Finally, both high-and low-quality samples are used to learn representations conducive to effective clustering. We build baselines on benchmark multimodal intent and dialogue act datasets. UMC shows remarkable improvements of 2-6% scores in clustering metrics over state-of-the-art methods, marking the first successful endeavor in this domain. The complete code and data are available at https://github.com/thuiar/UMC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly DataLiang Zhao, Shubin Ma, Bo Xu, Qingchen ZhangACM MM 2025 · 被引用 1 次
- Unsupervised Semantic Discovery via Global and Local Semantic Alignment in Multimodal ClusteringZhengzhong Zhu, Pei Zhou, Weihong Du, Shiquan Min 等AAAI 2026
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 被引用 737 次
相关 Paper
- New Intent Discovery with Pre-training and Contrastive LearningYuwei Zhang, Haode Zhang, Li-Ming Zhan, Xiao-Ming Wu 等ACL 2022 · 被引用 55 次
- Discovering New Intents with Deep Aligned ClusteringHanlei Zhang, Hua Xu, Ting-En Lin, Rui LyuAAAI 2021 · 被引用 138 次
- Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification with K-Means FeaturesTao Gui, Lizhi Qing, Qi Zhang, Jiacheng Ye 等AAAI 2020 · 被引用 2 次
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi 等AAAI 2026
- Deep Multiview Clustering by Contrasting Cluster AssignmentsJie Chen, Hua Mao, Wai Lok Woo, Xi PengICCV 2023 · 被引用 142 次
