Semantic-Enhanced Image Clustering
Shaotian Cai, Liping Qiu, Xiaojun Chen, Qin Zhang, Longteng Chen
摘要
Image clustering is an important and open challenging task in computer vision. Although many methods have been proposed to solve the image clustering task, they only explore images and uncover clusters according to the image features, thus being unable to distinguish visually similar but semantically different images. In this paper, we propose to investigate the task of image clustering with the help of visual-language pre-training model. Different from the zero-shot setting, in which the class names are known, we only know the number of clusters in this setting. Therefore, how to map images to a proper semantic space and how to cluster images from both image and semantic spaces are two key problems. To solve the above problems, we propose a novel image clustering method guided by the visual-language pre-training model CLIP, named Semantic-Enhanced Image Clustering (SIC). In this new method, we propose a method to map the given images to a proper semantic space first and efficient methods to generate pseudo-labels according to the relationships between images and semantics. Finally, we propose to perform clustering with consistency learning in both image space and semantic space, in a self-supervised learning fashion. The theoretical result of convergence analysis shows that our proposed method can converge at a sublinear speed. Theoretical analysis of expectation risk also shows that we can reduce the expectation risk by improving neighborhood consistency, increasing prediction confidence, or reducing neighborhood imbalance. Experimental results on five benchmark datasets clearly show the superiority of our new method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Image Clustering with External GuidanceYunfan Li, Peng Hu, Dezhong Peng, Jiancheng Lv 等ICML 2024 · 被引用 33 次
- Image Clustering Conditioned on Text CriteriaSehyun Kwon, Jaeseung Park, Minkyu Kim, Jaewoong Cho 等ICLR 2024 · 被引用 27 次
- Interactive Deep Clustering via Value MiningHonglin Liu, Peng Hu, Changqing Zhang, Yunfan Li 等NeurIPS 2024 · 被引用 24 次
- Mini-cluster Guided Long-tailed Deep ClusteringZhixin Li, Yuheng Jia, Guanliang Chen, Hui Liu 等ICLR 2026 · 被引用 11 次
- Multi-Level Cross-Modal Alignment for Image ClusteringLiping Qiu, Qin Zhang, Xiaojun Chen, Shaotian CaiAAAI 2024 · 被引用 8 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Self-Enhanced Image Clustering with Cross-Modal Semantic ConsistencyZihan Li, Wei Sun, Jing Hu, Jianhua Yin 等AAAI 2026
- iCLIP: Bridging Image Classification and Contrastive Language-Image Pre-training for Visual RecognitionYixuan Wei, Yue Cao, Zheng Zhang, Houwen Peng 等CVPR 2023
- Semantic-Augmented Image Clustering via Adaptive Multi-Modal CollaborationXiaohan Zhang, Chao Zhang, Deng Xu, Hong Yu 等AAAI 2026
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He 等ICML 2023 · 被引用 222 次
- S3A: Towards Realistic Zero-Shot Classification via Self Structural Semantic AlignmentSheng Zhang, Muzammal Naseer, Guangyi Chen, Zhiqiang Shen 等AAAI 2024 · 被引用 8 次
