Automatically Interpreting Millions of Features in Large Language Models
Gonçalo Paulo, Alex Mallen, Caden Juang, Nora Belrose
摘要
While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these activations into a higher-dimensional latent space which can be more easily interpretable. However, SAEs can have millions of distinct features, making it infeasible for humans to manually interpret each one. In this work, we build an open-source automated pipeline to generate and evaluate natural language interpretations for SAE features using LLMs. We test our framework on SAEs of varying sizes, activation functions, and losses, trained on two different open-weight LLMs. We introduce five new techniques to score the quality of interpretations that are cheaper to run than the previous state of the art. One of these techniques, intervention scoring, evaluates the interpretability of the effects of intervening on a feature, which we find explains features that are not recalled by existing methods. We propose guidelines for generating better interpretations that remain valid for a broader set of activating contexts, and discuss pitfalls with existing scoring techniques. Our code is available online.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 被引用 96 次
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersAdam Karvonen, James Chua, Clément Dumas, Kit Fraser-Taliente 等ICML 2026 · 被引用 42 次
- Automated Interpretability Metrics Do Not Distinguish Trained and Random TransformersThomas Heap, Tim Lawson, Lucy Farnik, Laurence AitchisonICLR 2026 · 被引用 32 次
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-TuningHelena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks 等ICML 2026 · 被引用 32 次
- I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse AutoencodersAndrey V. Galichin, Alexey Dontsov, Polina Druzhinina, Anton Razzhigaev 等AAAI 2026 · 被引用 31 次
它引用的顶会 Paper5
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- A Multimodal Automated Interpretability AgentTamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram 等ICML 2024 · 被引用 57 次
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh 等ICLR 2025 · 被引用 10 次
相关 Paper
- Evaluating SAE interpretability without generating explanationsGonçalo Paulo, Nora BelroseICLR 2026 · 被引用 2 次
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang 等ICML 2026
- Route Sparse Autoencoder to Interpret Large Language ModelsWei Shi, Sihang Li, Tao Liang, Mingyang Wan 等EMNLP 2025 · 被引用 1 次
- Interpretable and Steerable Concept Bottleneck Sparse AutoencodersAkshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 等CVPR 2026 · 被引用 6 次
- Sparse Autoencoders Do Not Find Canonical Units of AnalysisPatrick Leask, Bart Bussmann, Michael T. Pearce, Joseph Isaac Bloom 等ICLR 2025
