Ensembling Sparse Autoencoders
Soham Gadgil, Chris Lin, Su-In Lee
Abstract
Sparse autoencoders (SAEs) are used to decompose neural network activations into human-interpretable features. Typically, features learned by a single SAE are used for downstream applications. However, it has recently been shown that a single SAE captures only a limited subset of features that can be extracted from the activation space. Motivated by this limitation, we introduce and formalize SAE ensembles. Furthermore, we propose to ensemble multiple SAEs through naive bagging and boosting . In naive bagging, SAEs trained with different weight initializations are ensembled, whereas in boosting SAEs sequentially trained to minimize the residual error are ensembled. Theoretically, naive bagging and boosting are justified as approaches to reduce reconstruction error. Empirically, we evaluate our ensemble approaches with three settings of language models and SAE architectures. Our empirical results demonstrate that, compared to an expanded SAE that matches the number of features in the ensemble, ensembling SAEs improves the reconstruction of language model activations along with SAE stability. Additionally, on downstream tasks such as concept detection and spurious correlation removal, SAE ensembles achieve better performance, showing improved practical utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 202bacb4-2bf6-4aad-848c-0055301c32f9Builds on15
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar et al.NeurIPS 2025 · 168 citations
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène et al.NeurIPS 2023 · 125 citations
Related papers
- Are Sparse Autoencoders Useful? A Case Study in Sparse ProbingSubhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark et al.ICML 2025
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 96 citations
- Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept GeometrySai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, Demba BaNeurIPS 2025 · 65 citations
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 16 citations
- SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model InterpretabilityAdam Karvonen, Can Rager, Johnny Lin, Curt Tigges et al.ICML 2025
