MAGIC: Multi-Granularity Language-Informed Image Clustering
Xiaohan Zhang, Chao Zhang, Chunlin Chen, Huaxiong Li
Abstract
Image clustering is a fundamental unsupervised task in computer vision. Recent studies have explored incorporating external linguistic information to facilitate visual feature learning and thereby enhance clustering performance. Nevertheless, these methods typically rely on fixed lexical databases (e.g., WordNet) to generate language counterparts, leading to inter-modal semantic misalignment due to granularity discrepancy between visual and textual contents. Moreover, they often overlook the issue of intra-modal semantic redundancy caused by task-irrelevant knowledge. To address these challenges, we propose a new Multi-grAnularity lanGuage-informed Image Clustering method, dubbed MAGIC. To reduce semantic misalignment, we first prompt the vision-language models to generate multigranularity descriptions that capture rich image semantics, which are then integrated for effective multi-modal alignment. To alleviate semantic redundancy, we design information bottleneckinspired semantic adapters that adaptively refine and compress the semantically dense features into clustering-friendly representations under task guidance. A consensus representation is obtained by fusing the refined visual and textual features, which acts as a teacher to achieve cross-modal consistency and image clustering through a contrastive distillation framework. Extensive experiments on benchmarks demonstrate that MAGIC outperforms state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98644830-5218-4751-ab3d-7dbfc632611cBuilds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Semantic-Augmented Image Clustering via Adaptive Multi-Modal CollaborationXiaohan Zhang, Chao Zhang, Deng Xu, Hong Yu et al.AAAI 2026
- Image Clustering with External GuidanceYunfan Li, Peng Hu, Dezhong Peng, Jiancheng Lv et al.ICML 2024 · 33 citations
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- Boosting Medical Visual Understanding From Multi-Granular Language LearningZihan Li, Yiqing Wang, Sina Farsiu, Paul KinahanICLR 2026 · 6 citations
- Semantic-Enhanced Image ClusteringShaotian Cai, Liping Qiu, Xiaojun Chen, Qin Zhang et al.AAAI 2023 · 52 citations
