Residual Stream Analysis with Multi-Layer SAEs
Tim Lawson, Lucy Farnik, Conor J. Houghton, Laurence Aitchison
摘要
Sparse autoencoders (SAEs) are a promising approach to interpreting the internal representations of transformer language models. However, SAEs are usually trained separately on each transformer layer, making it difficult to use them to study how information flows across layers. To solve this problem, we introduce the multi-layer SAE (MLSAE): a single SAE trained on the residual stream activation vectors from every transformer layer. Given that the residual stream is understood to preserve information across layers, we expected MLSAE latents to 'switch on' at a token position and remain active at later layers. Interestingly, we find that individual latents are often active at a single layer for a given token or prompt, but the layer at which an individual latent is active may differ for different tokens or prompts. We quantify these phenomena by defining a distribution over layers and considering its variance. We find that the variance of the distributions of latent activations over layers is about two orders of magnitude greater when aggregating over tokens compared with a single token. For larger underlying models, the degree to which latents are active at multiple layers increases, which is consistent with the fact that the residual stream activation vectors at adjacent layers become more similar. Finally, we relax the assumption that the residual stream basis is the same at every layer by applying pre-trained tuned-lens transformations, but our findings remain qualitatively similar. Our results represent a new approach to understanding how representations change as they flow through transformers. We release our code to train and analyze MLSAEs at https://github.com/tim-lawson/mlsae.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Automated Interpretability Metrics Do Not Distinguish Trained and Random TransformersThomas Heap, Tim Lawson, Lucy Farnik, Laurence AitchisonICLR 2026 · 被引用 32 次
- Interpreting vision transformers via residual replacement modelJinyeong Kim, Junhyeok Kim, Yumin Shim, Joohyeok Kim 等NeurIPS 2025 · 被引用 4 次
- Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer GroupsDavide Ghilardi, Federico Belotti, Marco Molinari, Tao Ma 等EMNLP 2025
- Jacobian Sparse Autoencoders: Sparsify Computations, Not Just ActivationsLucy Farnik, Tim Lawson, Conor J. Houghton, Laurence AitchisonICML 2025
它引用的顶会 Paper15
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
相关 Paper
- Dense SAE Latents Are Features, Not BugsXiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu 等NeurIPS 2025 · 被引用 19 次
- LLM Layers Immediately Correct Each OtherArjun Patrawala, Jiahai Feng, Erik Jones, Jacob SteinhardtNeurIPS 2025 · 被引用 5 次
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 被引用 96 次
- Sparse autoencoders reveal selective remapping of visual concepts during adaptationHyesu Lim, Jinho Choi, Jaegul Choo, Steffen SchneiderICLR 2025
- Transferring Linear Features Across Language Models With Model StitchingAlan Chen, Jack Merullo, Alessandro Stolfo, Ellie PavlickNeurIPS 2025 · 被引用 17 次
