Image Clustering Conditioned on Text Criteria
Sehyun Kwon, Jaeseung Park, Minkyu Kim, Jaewoong Cho, Ernest K. Ryu, Kangwook Lee
Abstract
Classical clustering methods do not provide users with direct control of the clustering results, and the clustering results may not be consistent with the relevant criterion that a user has in mind. In this work, we present a new methodology for performing image clustering based on user-specified text criteria by leveraging modern vision-language models and large language models. We call our method Image Clustering Conditioned on Text Criteria (IC|TC), and it represents a different paradigm of image clustering. IC|TC requires a minimal and practical degree of human intervention and grants the user significant control over the clustering results in return. Our experiments show that IC|TC can effectively cluster images with various criteria, such as human action, physical location, or the person's mood, while significantly outperforming baselines. 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a01263f6-4e99-4062-8005-00410f2ba209Cited by top-tier papers6
- Conditional Representation Learning for Customized TasksHonglin Liu, Chao Sun, Peng Hu, Yunfan Li et al.NeurIPS 2025 · 6 citations
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual ConceptsJinho Choi, Hyesu Lim, Steffen Schneider, Jaegul ChooNeurIPS 2025 · 5 citations
- CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding SpaceSohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun OhCVPR 2026 · 2 citations
- Open Ad-hoc Categorization with Contextualized Feature LearningZilin Wang, Sangwoo Mo, Stella X. Yu, Sima Behpour et al.CVPR 2025
- Universal Guideline-Driven Image Clustering via a Hybrid LLM AgentWenliang Zhong, Rob Barton, Lucas Goncalves, Kushal Kumar et al.CVPR 2026
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- ESMC: MLLM-Based Embedding Selection for Explainable Multiple ClusteringXinyue Wang, Yuheng Jia, Hui Liu, Junhui HouAAAI 2026 · 1 citation
- Self-Enhanced Image Clustering with Cross-Modal Semantic ConsistencyZihan Li, Wei Sun, Jing Hu, Jianhua Yin et al.AAAI 2026
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi et al.AAAI 2026
- Multi-Modal Proxy Learning Towards Personalized Visual Multiple ClusteringJiawei Yao, Qi Qian, Juhua HuCVPR 2024 · 19 citations
- Semantic-Enhanced Image ClusteringShaotian Cai, Liping Qiu, Xiaojun Chen, Qin Zhang et al.AAAI 2023 · 52 citations
