Sparse Autoencoders Trained on the Same Data Learn Different Features
Gonçalo Paulo, Nora Belrose
摘要
Sparse autoencoders (SAEs) are a useful tool for uncovering human-interpretable features in the activations of large language models (LLMs). While some expect SAEs to find the true underlying features used by a model, our research shows that SAEs trained on the same model and data, differing only in the random seed used to initialize their weights, identify different sets of features. For example, in an SAE with 131K latents trained on a feedforward network in Llama 3 8B, only 30% of the features were shared across different seeds. We observed this phenomenon across multiple layers of three different LLMs, two datasets, and several SAE architectures. While ReLU SAEs trained with the L1 sparsity loss showed greater stability across seeds, SAEs using the state-of-the-art TopK activation function were more seed-dependent, even when controlling for the level of sparsity. Our results suggest that the set of features uncovered by an SAE should be viewed as a pragmatically useful decomposition of activation space, rather than an exhaustive and universal list of features ``truly used'' by the model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometrySai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba BaNeurIPS 2025 · 被引用 65 次
- Automated Interpretability Metrics Do Not Distinguish Trained and Random TransformersThomas Heap, Tim Lawson, Lucy Farnik, Laurence AitchisonICLR 2026 · 被引用 32 次
- Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski GeometryThomas Fel, Binxu Wang, Michael A. Lepori, Matthew Kowal 等ICLR 2026 · 被引用 28 次
- Base Models Know How to Reason, Thinking Models Learn WhenConstantin Venhoff, Iván Arcuschin, Phil Torr, Arthur Conmy 等ICML 2026 · 被引用 20 次
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper9
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar 等NeurIPS 2025 · 被引用 168 次
- Identifying Functionally Important Features with End-to-End Sparse Dictionary LearningDan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, Lee SharkeyNeurIPS 2024 · 被引用 81 次
- Git Re-Basin: Merging Models modulo Permutation SymmetriesSamuel K. Ainsworth, Jonathan Hayase, Siddhartha S. SrinivasaICLR 2023 · 被引用 32 次
相关 Paper
- Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse AutoencodersDavid Chanin, Adrià Garriga-AlonsoICML 2026 · 被引用 8 次
- Route Sparse Autoencoder to Interpret Large Language ModelsWei Shi, Sihang Li, Tao Liang, Mingyang Wan 等EMNLP 2025 · 被引用 1 次
- Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language ModelsPatrick Leask, Neel Nanda, Noura Al MoubayedICML 2025
- Are Sparse Autoencoders Useful? A Case Study in Sparse ProbingSubhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark 等ICML 2025
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 被引用 16 次
