LLMs Enable Bag-of-Texts Representations for Short-Text Clustering
I-Fan Lin, Faegheh Hasibi, Suzan Verberne
Abstract
In this paper, we propose a training-free method for unsupervised short text clustering that relies less on careful selection of embedders than other methods. In customer-facing chatbots, companies are dealing with large amounts of user utterances that need to be clustered according to their intent. In these settings, no labeled data is typically available, and the number of clusters is not known. Recent approaches to short-text clustering in label-free settings incorporate LLM output to refine existing embeddings. While LLMs can identify similar texts effectively, the resulting similarities may not be directly represented by distances in the dense vector space, as they depend on the original embedding. We therefore propose a method for transforming LLM judgments directly into a bag-of-texts representation in which texts are initialized to be equidistant, without assuming any prior distance relationships. Our method achieves comparable or superior results to state-of-the-art methods, but without embeddings optimization or assuming prior knowledge of clusters or labels. Experiments on diverse datasets and smaller LLMs show that our method is model agnostic and can be applied to any embedder, with relatively small LLMs, and different clustering methods. We also show how our method scales to large datasets, reducing the computational cost of the LLM use. The flexibility and scalability of our method make it more aligned with realworld training-free scenarios than existing clustering methods. Our source code is available here: https://anonymous.4open.science/ r/BoT_vector-2E0C/README.md
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 786f1b15-fb24-474c-8540-b29cc29ed901Builds on5
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
- ClusterLLM: Large Language Models as a Guide for Text ClusteringYuwei Zhang, Zihan Wang, Jingbo ShangEMNLP 2023 · 43 citations
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen et al.ACL 2020 · 16 citations
- TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data ConsistencyHenry Peng Zou, Zhengyao Gu, Yue Zhou, Yankai Chen et al.ACL 2025 · 3 citations
Related papers
- Improving Clustering with Positive Pairs Generated from LLM-Driven LabelsXiaotong Zhang, Ying LiEMNLP 2025
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- Beyond prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering RepresentationsYu Fei, Zhao Meng, Ping Nie, Roger Wattenhofer et al.EMNLP 2022 · 13 citations
- PromptBoosting: Black-Box Text Classification with Ten Forward PassesBairu Hou, Joe O'Connor, Jacob Andreas, Shiyu Chang et al.ICML 2023 · 56 citations
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual RepresentationsJohn M. Giorgi, Osvald Nitski, Bo Wang, Gary D. BaderACL 2021
