Sparse Autoencoders for Hypothesis Generation
Rajiv Movva, Kenny Peng, Nikhil Garg, Jon M. Kleinberg, Emma Pierson
Abstract
We describe HYPOTHESAES, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HYPOTHESAES has three steps: (1) train a sparse autoencoder on text embeddings to produce interpretable features describing the data distribution, (2) select features that predict the target variable, and (3) generate a natural language interpretation of each feature (e.g., mentions being surprised or shocked) using an LLM. Each interpretation serves as a hypothesis about what predicts the target variable. Compared to baselines, our method better identifies reference hypotheses on synthetic datasets (at least +0.06 in F1) and produces more predictive hypotheses on real datasets (∼twice as many significant findings), despite requiring 1-2 orders of magnitude less compute than recent LLM-based methods. HYPOTHESAES also produces novel discoveries on two well-studied tasks: explaining partisan differences in Congressional speeches and identifying drivers of engagement with online headlines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ce3c793-1b69-4f33-9178-82215a167918Cited by top-tier papers12
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-TuningHelena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks et al.ICML 2026 · 32 citations
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 27 citations
- BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language ModelsLindia Tjuatja, Graham NeubigACL 2025 · 5 citations
- Exploratory Causal Inference in SAEnceTommaso Mencattini, Riccardo Cadei, Francesco LocatelloICLR 2026 · 4 citations
- Can SAEs reveal and mitigate racial biases of LLMs in healthcare?Hiba Ahsan, Byron C. WallaceICLR 2026 · 1 citation
Builds on8
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong et al.ICLR 2021 · 357 citations
- Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooMMichelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer et al.CHI 2024 · 46 citations
- Explaining Datasets in Words: Statistical Models with Natural Language ParametersRuiqi Zhong, Heng Wang, Dan Klein, Jacob SteinhardtNeurIPS 2024 · 26 citations
Related papers
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 16 citations
- Compute Optimal Inference and Provable Amortisation Gap in Sparse AutoencodersCharles O'Neill, Alim Gumran, David A. KlindtICML 2025
- Interpretable Embeddings with Sparse Autoencoders: A Data Analysis ToolkitNick Jiang, Xiaoqing Sun, Lisa Dunlap, Lewis Smith et al.ICML 2026
- Step-Level Sparse Autoencoder for Reasoning Process InterpretationXuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu et al.ICML 2026 · 2 citations
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
