ClusterLLM: Large Language Models as a Guide for Text Clustering
Yuwei Zhang, Zihan Wang, Jingbo Shang
Abstract
We introduce CLUSTERLLM, a novel text clustering framework that leverages feedback from an instruction-tuned large language model, such as ChatGPT. Compared with traditional unsupervised methods that builds upon "small" embedders, CLUSTERLLM exhibits two intriguing advantages: (1) it enjoys the emergent capability of LLM even if its embeddings are inaccessible; and (2) it understands the user's preference on clustering through textual instruction and/or a few annotated data. First, we prompt ChatGPT for insights on clustering perspective by constructing hard triplet questions <does A better correspond to B than C>, where A, B and C are similar data points that belong to different clusters according to small embedder. We empirically show that this strategy is both effective for fine-tuning small embedder and cost-efficient to query ChatGPT. Second, we prompt ChatGPT for helps on clustering granularity by carefully designed pairwise questions <do A and B belong to the same category>, and tune the granularity from cluster hierarchies that is the most consistent with the ChatGPT answers. Extensive experiments on 14 datasets show that CLUSTERLLM consistently improves clustering quality, at an average cost of ∼$0.6 1 per dataset. The code will be available at https: //github.com/zhang-yu-wei/ClusterLLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27b2f570-9e6b-43af-afea-b956a01f04e0Cited by top-tier papers27
- AutoDiscovery: Open-ended Scientific Discovery via Bayesian SurpriseDhruv Agarwal, Bodhisattwa Prasad Majumder, Reece Adamson, Megha Chakravorty et al.NeurIPS 2025 · 35 citations
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-ExplorationChengbo Zheng, Yuanhao Zhang, Zeyu Huang, Chuhan Shi et al.UIST 2024 · 19 citations
- Summaries as Centroids for Interpretable and Scalable Text ClusteringJairo Diaz RodriguezICLR 2026 · 11 citations
- Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect ClusteringKun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang et al.EMNLP 2025 · 8 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Open-World Semi-Supervised LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2022 · 246 citations
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
Related papers
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- Answer is All You Need: Instruction-following Text Embedding via Answering the QuestionLetian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa et al.ACL 2024
- Optimized Algorithms for Text Clustering with LLM-Generated ConstraintsChaoqi Jia, Weihong Wu, Longkun Guo, Zhigang Lu et al.AAAI 2026
- CEIL: A General Classification-Enhanced Iterative Learning Framework for Text ClusteringMingjun Zhao, Mengzhen Wang, Yinglong Ma, Di Niu et al.WWW 2023 · 2 citations
- LLMs Enable Bag-of-Texts Representations for Short-Text ClusteringI-Fan Lin, Faegheh Hasibi, Suzan VerberneACL 2026 · 2 citations
