Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
Davide Ghilardi, Federico Belotti, Marco Molinari, Tao Ma, Matteo Palmonari
Abstract
Sparse Autoencoders (SAEs) have recently been employed as a promising unsupervised approach for understanding the representations of layers of Large Language Models (LLMs). However, with the growth in model size and complexity, training SAEs is computationally intensive, as typically one SAE is trained for each model layer. To address such limitation, we propose Group-SAE, a novel strategy to train SAEs. Our method considers the similarity of the residual stream representations between contiguous layers to group similar layers and train a single SAE per group. To balance the trade-off between efficiency and performance, we further introduce AMAD (Average Maximum Angular Distance), an empirical metric that guides the selection of an optimal number of groups based on representational similarity across layers. Experiments on models from the Pythia family show that our approach significantly accelerates training with minimal impact on reconstruction quality and comparable downstream task performance and interpretability over baseline SAEs trained layer by layer. This method provides an efficient and scalable strategy for training SAEs in modern LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36ac7de0-93fd-427c-acdf-989e828debb6Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LNPengxiang Li, Lu Yin, Shiwei LiuICLR 2025
- The Unreasonable Ineffectiveness of the Deeper LayersAndrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso et al.ICLR 2025
Related papers
- Residual Stream Analysis with Multi-Layer SAEsTim Lawson, Lucy Farnik, Conor J. Houghton, Laurence AitchisonICLR 2025
- Jacobian Sparse Autoencoders: Sparsify Computations, Not Just ActivationsLucy Farnik, Tim Lawson, Conor J. Houghton, Laurence AitchisonICML 2025
- Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language ModelsPatrick Leask, Neel Nanda, Noura Al MoubayedICML 2025
- Dense SAE Latents Are Features, Not BugsXiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu et al.NeurIPS 2025 · 19 citations
- Low-Rank Adapting Models for Sparse AutoencodersMatthew Chen, Joshua Engels, Max TegmarkICML 2025
