Goal-Driven Explainable Clustering via Language Descriptions
Zihan Wang, Jingbo Shang, Ruiqi Zhong
Abstract
Unsupervised clustering is widely used to explore large corpora, but existing formulations neither consider the users' goals nor explain clusters' meanings. We propose a new task formulation, "Goal-Driven Clustering with Explanations" (GoalEx), which represents both the goal and the explanations as free-form language descriptions. For example, to categorize the errors made by a summarization system, the input to GoalEx is a corpus of annotator-written comments for system-generated summaries and a goal description "cluster the comments based on why the annotators think the summary is imperfect."; the outputs are text clusters each with an explanation ("this cluster mentions that the summary misses important context information."), which relates to the goal and accurately explains which comments should (not) belong to a cluster. To tackle GoalEx, we prompt a language model with "[corpus subset] + [goal] + Brainstorm a list of explanations each representing a cluster."; then we classify whether each sample belongs to a cluster based on its explanation; finally, we use integer linear programming to select a subset of candidate clusters to cover most samples while minimizing overlaps. Under both automatic and human evaluation on corpora with or without labels, our method produces more accurate and goal-related explanations than prior methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2b36238-dcb0-4d5a-9ea8-afd462781ae0Cited by top-tier papers16
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn et al.NeurIPS 2023 · 81 citations
- Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooMMichelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer et al.CHI 2024 · 46 citations
- ClusterLLM: Large Language Models as a Guide for Text ClusteringYuwei Zhang, Zihan Wang, Jingbo ShangEMNLP 2023 · 43 citations
- Explaining Datasets in Words: Statistical Models with Natural Language ParametersRuiqi Zhong, Heng Wang, Dan Klein, Jacob SteinhardtNeurIPS 2024 · 26 citations
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu et al.IEEE VIS 2024 · 23 citations
Builds on11
- Open-World Semi-Supervised LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2022 · 246 citations
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck et al.ICLR 2022 · 178 citations
- Scaling Laws for Generative Mixed-Modal Language ModelsArmen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu et al.ICML 2023 · 149 citations
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 121 citations
Related papers
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi et al.AAAI 2026
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- SummAct: Uncovering User Intentions Through Interactive Behaviour SummarisationGuanhua Zhang, Mohamed Adel Naguib Ahmed, Zhiming Hu, Andreas BullingCHI 2025 · 5 citations
- Cluster Explanation via Polyhedral DescriptionsConnor Lawless, Oktay GünlükICML 2023 · 14 citations
- Latent Principle Discovery for Language Model Self-ImprovementKeshav Ramji, Tahira Naseem, Ramón Fernandez AstudilloNeurIPS 2025 · 2 citations
