Goal-Driven Explainable Clustering via Language Descriptions
Zihan Wang, Jingbo Shang, Ruiqi Zhong
摘要
Unsupervised clustering is widely used to explore large corpora, but existing formulations neither consider the users' goals nor explain clusters' meanings. We propose a new task formulation, "Goal-Driven Clustering with Explanations" (GoalEx), which represents both the goal and the explanations as free-form language descriptions. For example, to categorize the errors made by a summarization system, the input to GoalEx is a corpus of annotator-written comments for system-generated summaries and a goal description "cluster the comments based on why the annotators think the summary is imperfect."; the outputs are text clusters each with an explanation ("this cluster mentions that the summary misses important context information."), which relates to the goal and accurately explains which comments should (not) belong to a cluster. To tackle GoalEx, we prompt a language model with "[corpus subset] + [goal] + Brainstorm a list of explanations each representing a cluster."; then we classify whether each sample belongs to a cluster based on its explanation; finally, we use integer linear programming to select a subset of candidate clusters to cover most samples while minimizing overlaps. Under both automatic and human evaluation on corpora with or without labels, our method produces more accurate and goal-related explanations than prior methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn 等NeurIPS 2023 · 被引用 81 次
- Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooMMichelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer 等CHI 2024 · 被引用 46 次
- ClusterLLM: Large Language Models as a Guide for Text ClusteringYuwei Zhang, Zihan Wang, Jingbo ShangEMNLP 2023 · 被引用 43 次
- Explaining Datasets in Words: Statistical Models with Natural Language ParametersRuiqi Zhong, Heng Wang, Dan Klein, Jacob SteinhardtNeurIPS 2024 · 被引用 26 次
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu 等IEEE VIS 2024 · 被引用 23 次
它引用的顶会 Paper11
- Open-World Semi-Supervised LearningKaidi Cao, Maria Brbic, Jure LeskovecICLR 2022 · 被引用 246 次
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 被引用 194 次
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck 等ICLR 2022 · 被引用 178 次
- Scaling Laws for Generative Mixed-Modal Language ModelsArmen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu 等ICML 2023 · 被引用 149 次
- Contextualized Weak Supervision for Text ClassificationDheeraj Mekala, Jingbo ShangACL 2020 · 被引用 121 次
相关 Paper
- MLLM Enriched Explainable Multiple ClusteringShan Zhang, Liangrui Ren, Qiaoyu Tan, Carlotta Domeniconi 等AAAI 2026
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- SummAct: Uncovering User Intentions Through Interactive Behaviour SummarisationGuanhua Zhang, Mohamed Adel Naguib Ahmed, Zhiming Hu, Andreas BullingCHI 2025 · 被引用 5 次
- Cluster Explanation via Polyhedral DescriptionsConnor Lawless, Oktay GünlükICML 2023 · 被引用 14 次
- Latent Principle Discovery for Language Model Self-ImprovementKeshav Ramji, Tahira Naseem, Ramón Fernandez AstudilloNeurIPS 2025 · 被引用 2 次
