Explaining Datasets in Words: Statistical Models with Natural Language Parameters
Ruiqi Zhong, Heng Wang, Dan Klein, Jacob Steinhardt
Abstract
To make sense of massive data, we often fit simplified models and then interpret the parameters; for example, we cluster the text embeddings and then interpret the mean parameters of each cluster. However, these parameters are often high-dimensional and hard to interpret. To make model parameters directly interpretable, we introduce a family of statistical models -- including clustering, time series, and classification models -- parameterized by natural language predicates. For example, a cluster of text about COVID could be parameterized by the predicate"discusses COVID". To learn these statistical models effectively, we develop a model-agnostic algorithm that optimizes continuous relaxations of predicate parameters with gradient descent and discretizes them by prompting language models (LMs). Finally, we apply our framework to a wide range of problems: taxonomizing user chat dialogues, characterizing how they evolve across time, finding categories where one language model is better than the other, clustering math problems based on subareas, and explaining visual features in memorable images. Our framework is highly versatile, applicable to both textual and visual domains, can be easily steered to focus on specific properties (e.g. subareas), and explains sophisticated concepts that classical methods (e.g. n-gram analysis) struggle to produce.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3534f6a-91e2-4e5a-864a-8d4def5e0922Cited by top-tier papers4
- Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of ImagesBoyang Deng, Songyou Peng, Kyle Genova, Gordon Wetzstein et al.ICCV 2025 · 4 citations
- ProxAnn: Use-Oriented Evaluations of Topic Models and Document ClusteringAlexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan Lee Boyd-Graber, Philip ResnikACL 2025
- Sparse Autoencoders for Hypothesis GenerationRajiv Movva, Kenny Peng, Nikhil Garg, Jon M. Kleinberg et al.ICML 2025
- RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language ModelsYang Yang, Hua XU, Zhangyi Hu, Yutao YueICML 2026
Builds on17
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
Related papers
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Basel Mousi, Nadir Durrani, Fahim DalviEMNLP 2023 · 1 citation
- Automated Statistical Model Discovery with Language ModelsMichael Y. Li, Emily B. Fox, Noah D. GoodmanICML 2024 · 36 citations
- LLM Processes: Numerical Predictive Distributions Conditioned on Natural LanguageJames Requeima, John Bronskill, Dami Choi, Richard E. Turner et al.NeurIPS 2024 · 72 citations
- Customized Multiple Clustering via Multi-Modal Subspace Proxy LearningJiawei Yao, Qi Qian, Juhua HuNeurIPS 2024 · 17 citations
- Hi-Time: Hierarchical Latent Prediction for Multivariate Time Series ClassificationKun Zeng, Wu Binquan, Qianli MaICML 2026
