Hierarchical Semantic Alignment for Image Clustering
Xingyu Zhu, Beier Zhu, Yunfan Li, Junfeng Fang, Shuo Wang, Kesen Zhao, Hanwang Zhang
Abstract
Image clustering is a classic problem in computer vision, which categorizes images into different groups. Recent studies utilize nouns as external semantic knowledge to improve clustering performance. However, these methods often overlook the inherent ambiguity of nouns, which can distort semantic representations and degrade clustering quality. To address this issue, we propose a hierarChical semAntic alignmEnt method for image clustering, dubbed CAE, which improves clustering performance in a training-free manner. In our approach, we incorporate two complementary types of textual semantics: caption-level descriptions, which convey fine-grained attributes of image content, and noun-level concepts, which represent high-level object categories. We first select relevant nouns from WordNet and descriptions from caption datasets to construct a semantic space aligned with image features. Then, we align image features with selected nouns and captions via optimal transport to obtain a more discriminative semantic space. Finally, we combine the enhanced semantic and image features to perform clustering. Extensive experiments across 8 datasets demonstrate the effectiveness of our method, notably surpassing the state-of-the-art training-free approach with a 4.2% improvement in accuracy and a 2.9% improvement in adjusted rand index (ARI) on the ImageNet-1K dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e5c3140-dfb0-47a5-bc03-665731524d6dCited by top-tier papers6
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- Real-Time Motion-Controllable Autoregressive Video DiffusionKesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou et al.ICLR 2026 · 10 citations
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang et al.ICLR 2026 · 9 citations
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!Junbao Zhou, Yuan Zhou, Kesen Zhao, Qingshan Xu et al.ICLR 2026 · 7 citations
- GuardAlign: Test-time Safety Alignment in Multimodal Large Language ModelsXingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang et al.ICLR 2026 · 2 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Contrastive ClusteringYunfan Li, Peng Hu, Jerry Zitao Liu, Dezhong Peng et al.AAAI 2021 · 798 citations
- Deep Comprehensive Correlation Mining for Image ClusteringJianlong Wu, Keyu Long, Fei Wang, Chen Qian et al.ICCV 2019 · 191 citations
Related papers
- Image Clustering with External GuidanceYunfan Li, Peng Hu, Dezhong Peng, Jiancheng Lv et al.ICML 2024 · 33 citations
- Multi-Level Cross-Modal Alignment for Image ClusteringLiping Qiu, Qin Zhang, Xiaojun Chen, Shaotian CaiAAAI 2024 · 8 citations
- Noise-aware Learning from Web-crawled Image-Text Data for Image CaptioningWooyoung Kang, Jonghwan Mun, Sungjun Lee, Byungseok RohICCV 2023 · 33 citations
- Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingXuanpu Zhao, Dianmo Sheng, Zhentao Tan, Zhiwei Zhao et al.AAAI 2025 · 2 citations
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen et al.CVPR 2024 · 38 citations
