Sparse Autoencoders for Hypothesis Generation
Rajiv Movva, Kenny Peng, Nikhil Garg, Jon M. Kleinberg, Emma Pierson
摘要
We describe HYPOTHESAES, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HYPOTHESAES has three steps: (1) train a sparse autoencoder on text embeddings to produce interpretable features describing the data distribution, (2) select features that predict the target variable, and (3) generate a natural language interpretation of each feature (e.g., mentions being surprised or shocked) using an LLM. Each interpretation serves as a hypothesis about what predicts the target variable. Compared to baselines, our method better identifies reference hypotheses on synthetic datasets (at least +0.06 in F1) and produces more predictive hypotheses on real datasets (∼twice as many significant findings), despite requiring 1-2 orders of magnitude less compute than recent LLM-based methods. HYPOTHESAES also produces novel discoveries on two well-studied tasks: explaining partisan differences in Congressional speeches and identifying drivers of engagement with online headlines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-TuningHelena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks 等ICML 2026 · 被引用 32 次
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference DataRajiv Movva, Smitha Milli, Sewon Min, Emma PiersonICLR 2026 · 被引用 27 次
- BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language ModelsLindia Tjuatja, Graham NeubigACL 2025 · 被引用 5 次
- Exploratory Causal Inference in SAEnceTommaso Mencattini, Riccardo Cadei, Francesco LocatelloICLR 2026 · 被引用 4 次
- Can SAEs reveal and mitigate racial biases of LLMs in healthcare?Hiba Ahsan, Byron C. WallaceICLR 2026 · 被引用 1 次
它引用的顶会 Paper8
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
- Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooMMichelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer 等CHI 2024 · 被引用 46 次
- Explaining Datasets in Words: Statistical Models with Natural Language ParametersRuiqi Zhong, Heng Wang, Dan Klein, Jacob SteinhardtNeurIPS 2024 · 被引用 26 次
相关 Paper
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 被引用 16 次
- Compute Optimal Inference and Provable Amortisation Gap in Sparse AutoencodersCharles O'Neill, Alim Gumran, David A. KlindtICML 2025
- Interpretable Embeddings with Sparse Autoencoders: A Data Analysis ToolkitNick Jiang, Xiaoqing Sun, Lisa Dunlap, Lewis Smith 等ICML 2026
- Step-Level Sparse Autoencoder for Reasoning Process InterpretationXuan Yang, Jiayu Liu, Yuhang Lai, Hao Xu 等ICML 2026 · 被引用 2 次
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello 等NeurIPS 2024 · 被引用 26 次
