Enhancing Automated Interpretability with Output-Centric Feature Descriptions
Yoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger, Mor Geva
Abstract
Automated interpretability pipelines generate natural language descriptions for the concepts represented by features in large language models (LLMs), such as plants or the first word in a sentence. These descriptions are derived using inputs that activate the feature, which may be a dimension or a direction in the model's representation space. However, identifying activating inputs is costly, and the mechanistic role of a feature in model behavior is determined both by how inputs cause a feature to activate and by how feature activation affects outputs. Using steering evaluations, we reveal that current pipelines provide descriptions that fail to capture the causal effect of the feature on outputs. To fix this, we propose efficient, output-centric methods for automatically generating feature descriptions. These methods use the tokens weighted higher after feature stimulation or the highest weight tokens after applying the vocabulary "unembedding" head directly to the feature. Our output-centric descriptions better capture the causal effect of a feature on model outputs than input-centric descriptions, but combining the two leads to the best performance on both input and output evaluations. Lastly, we show that output-centric descriptions can be used to find inputs that activate features previously thought to be "dead".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer et al.NeurIPS 2025 · 12 citations
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez et al.EMNLP 2025 · 10 citations
- Manipulating Feature Visualizations with Gradient SlingshotsDilyara Bareeva, Marina M.-C. Höhne, Alexander Warnecke, Lukas Pirch et al.NeurIPS 2025 · 8 citations
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 8 citations
- Circuit Insights: Towards Interpretability Beyond ActivationsElena Golimblevskaia, Aakriti Jain, Bruno Puri, Ammar Ibrahim et al.ICLR 2026 · 4 citations
Builds on16
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- A Multimodal Automated Interpretability AgentTamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram et al.ICML 2024 · 57 citations
Related papers
- Language Models Can Explain Visual Features via SteeringJavier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos et al.CVPR 2026 · 2 citations
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured LanguageAngie W. Boggust, Donghao Ren, Yannick Assogba, Dominik Moritz et al.ICLR 2026 · 3 citations
- Automatically Interpreting Millions of Features in Large Language ModelsGonçalo Paulo, Alex Mallen, Caden Juang, Nora BelroseICML 2025
- Constructing Interpretable Features from Compositional Neuron GroupsOr David Shafran, Atticus Geiger, Mor GevaACL 2026 · 4 citations
- SAEs Are Good for Steering - If You Select the Right FeaturesDana Arad, Aaron Mueller, Yonatan BelinkovEMNLP 2025
