Explaining Datasets in Words: Statistical Models with Natural Language Parameters
Ruiqi Zhong, Heng Wang, Dan Klein, Jacob Steinhardt
摘要
To make sense of massive data, we often fit simplified models and then interpret the parameters; for example, we cluster the text embeddings and then interpret the mean parameters of each cluster. However, these parameters are often high-dimensional and hard to interpret. To make model parameters directly interpretable, we introduce a family of statistical models -- including clustering, time series, and classification models -- parameterized by natural language predicates. For example, a cluster of text about COVID could be parameterized by the predicate"discusses COVID". To learn these statistical models effectively, we develop a model-agnostic algorithm that optimizes continuous relaxations of predicate parameters with gradient descent and discretizes them by prompting language models (LMs). Finally, we apply our framework to a wide range of problems: taxonomizing user chat dialogues, characterizing how they evolve across time, finding categories where one language model is better than the other, clustering math problems based on subareas, and explaining visual features in memorable images. Our framework is highly versatile, applicable to both textual and visual domains, can be easily steered to focus on specific properties (e.g. subareas), and explains sophisticated concepts that classical methods (e.g. n-gram analysis) struggle to produce.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of ImagesBoyang Deng, Songyou Peng, Kyle Genova, Gordon Wetzstein 等ICCV 2025 · 被引用 4 次
- ProxAnn: Use-Oriented Evaluations of Topic Models and Document ClusteringAlexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan Lee Boyd-Graber, Philip ResnikACL 2025
- Sparse Autoencoders for Hypothesis GenerationRajiv Movva, Kenny Peng, Nikhil Garg, Jon M. Kleinberg 等ICML 2025
- RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language ModelsYang Yang, Hua XU, Zhangyi Hu, Yutao YueICML 2026
它引用的顶会 Paper17
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 被引用 340 次
相关 Paper
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Basel Mousi, Nadir Durrani, Fahim DalviEMNLP 2023 · 被引用 1 次
- Automated Statistical Model Discovery with Language ModelsMichael Y. Li, Emily B. Fox, Noah D. GoodmanICML 2024 · 被引用 36 次
- LLM Processes: Numerical Predictive Distributions Conditioned on Natural LanguageJames Requeima, John Bronskill, Dami Choi, Richard E. Turner 等NeurIPS 2024 · 被引用 72 次
- Customized Multiple Clustering via Multi-Modal Subspace Proxy LearningJiawei Yao, Qi Qian, Juhua HuNeurIPS 2024 · 被引用 17 次
- Hi-Time: Hierarchical Latent Prediction for Multivariate Time Series ClassificationKun Zeng, Wu Binquan, Qianli MaICML 2026
