Improving Clustering with Positive Pairs Generated from LLM-Driven Labels
Xiaotong Zhang, Ying Li
Abstract
Traditional unsupervised clustering methods, which often rely on contrastive training of embedders, suffer from a lack of label knowledge, resulting in suboptimal performance. Furthermore, the presence of potential false negatives can destabilize the training process. Hence, we propose to improve clustering with Positive Pairs generated from LLM-driven Labels (PPLL). In the proposed framework, LLM is initially employed to cluster the data and generate corresponding mini-cluster labels. Subsequently, positive pairs are constructed based on these labels, and an embedder is trained using BYOL to obviate the need for negative pairs. Following training, the acquired label knowledge is integrated into K-means clustering. This framework enables the integration of label information throughout the training and inference processes, while mitigating the reliance on negative pairs. Additionally, it generates interpretable labels for improved understanding of clustering results. Empirical evaluations on a range of datasets demonstrate that our proposed framework consistently surpasses state-of-the-art baselines, achieving superior performance, robustness, and computational efficiency for diverse text clustering applications. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 014bb413-7b02-45c2-b44d-faa5af334a71Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Prototypical Contrastive Learning of Unsupervised RepresentationsJunnan Li, Pan Zhou, Caiming Xiong, Steven C. H. HoiICLR 2021 · 484 citations
- Graph Contrastive ClusteringHuasong Zhong, Jianlong Wu, Chong Chen, Jianqiang Huang et al.ICCV 2021 · 163 citations
Related papers
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- ClusterLLM: Large Language Models as a Guide for Text ClusteringYuwei Zhang, Zihan Wang, Jingbo ShangEMNLP 2023 · 43 citations
- LLMs Enable Bag-of-Texts Representations for Short-Text ClusteringI-Fan Lin, Faegheh Hasibi, Suzan VerberneACL 2026 · 2 citations
- When Phrases Meet Probabilities: Enabling Open Relation Extraction with Cooperating Large Language ModelsJiaxin Wang, Lingling Zhang, Wee Sun Lee, Yujie Zhong et al.ACL 2024
- Optimized Algorithms for Text Clustering with LLM-Generated ConstraintsChaoqi Jia, Weihong Wu, Longkun Guo, Zhigang Lu et al.AAAI 2026
