Summaries as Centroids for Interpretable and Scalable Text Clustering
Jairo Diaz Rodriguez
摘要
We introduce k-NLPmeans and k-LLMmeans, text-clustering variants of k-means that periodically replace numeric centroids with textual summaries. The key idea, summary-as-centroid, retains k-means assignments in embedding space while producing human-readable, auditable cluster prototypes. The method is LLM-optional: k-NLPmeans uses lightweight, deterministic summarizers, enabling offline, low-cost, and stable operation; k-LLMmeans is a drop-in upgrade that uses an LLM for summaries under a fixed per-iteration budget whose cost does not grow with dataset size. We also present a mini-batch extension for real-time clustering of streaming text. Across diverse datasets, embedding models, and summarization strategies, our approach consistently outperforms classical baselines and approaches the accuracy of recent LLM-based clustering-without extensive LLM calls. Finally, we provide a case study on sequential text streams and release a StackExchange-derived benchmark for evaluating streaming text clustering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Unsupervised Story Discovery from Continuous News Streams via Scalable Thematic EmbeddingSusik Yoon, Dongha Lee, Yunyi Zhang, Jiawei HanSIGIR 2023 · 被引用 8 次
- PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets StreamSusik Yoon, Hou Pong Chan, Jiawei HanWWW 2023 · 被引用 13 次
- LLMs Enable Bag-of-Texts Representations for Short-Text ClusteringI-Fan Lin, Faegheh Hasibi, Suzan VerberneACL 2026 · 被引用 2 次
- No Need to Talk: Asynchronous Mixture of Language ModelsAnastasiia Filippova, Angelos Katharopoulos, David Grangier, Ronan CollobertICLR 2025
- Goal-Driven Explainable Clustering via Language DescriptionsZihan Wang, Jingbo Shang, Ruiqi ZhongEMNLP 2023 · 被引用 19 次
